Senior Network Engineer — AI Datacenter

Nava

Bengaluru

On-site

INR 1,800,000 - 2,400,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nava designs, deploys, and operates large-scale GPU infrastructure for AI inference at scale. The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, delivering lossless, high-throughput networking for GPU-to-GPU traffic and distributed training workloads.

We seek hands-on engineers capable of building leaf-spine fabrics, automating provisioning, and leading deployment of GPU clusters.

Qualifications

  • 5–8 years in datacenter or production network engineering.
  • Hands-on experience with BGP, EVPN-VXLAN, and RDMA fabrics.
  • Strong knowledge of core protocols TCP/IP, IPv4/IPv6, DNS, DHCP.
  • Automation skills in Python, Ansible, and Terraform.
  • Comfortable in fast-paced on-call production environments.

Responsibilities

  • Design and deliver leaf-spine fabrics using BGP and EVPN-VXLAN, and implement RoCEv2 and InfiniBand lossless networking for GPU backend traffic.
  • Build and extend network automation in Python, Ansible, and Terraform to provision, validate, and operate the fabric.
  • Lead deployment and turn-up of new GPU clusters, including acceptance testing and performance validation against line-rate targets.
  • Triage and resolve production network incidents; tune congestion control to keep RDMA traffic lossless under real load.
  • Partner with Compute and Storage engineers to integrate the network end-to-end and eliminate bottlenecks.
  • Contribute to standards and mentor Network Engineers.

Skills

BGP
EVPN-VXLAN
RoCEv2
InfiniBand
Python
Ansible
Terraform
TCP/IP
Leaf-spine

Tools

iproute2
netlink

Job description

About Nava

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.

About The Team

The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, from architecture and design through deployment, automation, and production operation. We build the lossless, high-throughput networks that carry GPU-to-GPU traffic for distributed training and low-latency inference. This is a hands-on team where the network is treated as code.

Responsibilities
  • Design and deliver leaf-spine fabrics using BGP and EVPN-VXLAN, and implement RoCEv2 and InfiniBand lossless networking for GPU backend traffic.
  • Build and extend network automation in Python, Ansible, and Terraform to provision, validate, and operate the fabric.
  • Lead deployment and turn-up of new GPU clusters, including acceptance testing and performance validation against line-rate targets.
  • Triage and resolve production network incidents; tune congestion control (PFC/ECN) to keep RDMA traffic lossless under real load.
  • Partner with Compute and Storage engineers to integrate the network end-to-end and eliminate bottlenecks.
  • Contribute to standards and mentor Network Engineers.
Required Qualifications
  • 5-8 years in datacenter or production network engineering with strong hands-on BGP and EVPN-VXLAN experience.
  • Strong hands-on command of core networking protocols: BGP, OSPF, IS-IS, TCP/IP, IPv4 and IPv6, DNS, DHCP, and MPLS.
  • Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL/TLS.
  • Solid understanding of datacenter fabric concepts (leaf-spine, EVPN-VXLAN) and RDMA fabrics (RoCEv2 and InfiniBand), including lossless Ethernet.
  • Automation skills: Python plus Ansible and/or Terraform.
  • Comfortable operating in a fast-paced, on-call production environment.
Preferred Qualifications
  • Experience with NVIDIA networking (Spectrum-X, Quantum InfiniBand, BlueField DPUs) and NCCL traffic patterns.
  • Experience operating GPU clusters for large-scale distributed training or inference.
  • Familiarity with network telemetry, streaming analytics, and closed-loop automation.
  • Relevant certifications (e.g., CCNP or vendor equivalents).
  • Experience with Cisco and/or Arista platforms is a strong plus.
Technology environment

At Nava, Our Network Infrastructure Is Deeply Integrated With Modern AI Workloads - designed For Scale, Performance, And Automation. Below Is An Overview Of The Key Technologies And Platforms You’ll Work With Daily

  • Hardware & Interconnects: NVIDIA DGX/HGX systems (including Blackwell architecture), NVLink, Spectrum and Quantum-based switches, BlueField DPUs, and InfiniBand fabrics.
  • RDMA & Lossless Networking: RoCEv2 and InfiniBand RDMA, with deep familiarity in congestion control mechanisms (PFC, ECN) and lossless Ethernet configuration.
  • Control & Data Plane Protocols: BGP (including MP-BGP for EVPN), OSPF, IS-IS, EVPN-VXLAN for fabric virtualization, MPLS, and full IPv4/IPv6 stack support.
  • Automation & Tooling: Python for custom tooling, Ansible for configuration management, Terraform for infrastructure-as-code, and Linux-based toolchains (e.g., iproute2, netlink).
  • Observability & Operations: Telemetry via streaming telemetry (gNMI/sFlow), integration with monitoring stacks (Prometheus/Grafana), and experience with closed-loop automation for anomaly detection and remediation.

You’ll operate in a unified stack where networking, compute, and storage are co-designed - making deep technical understanding and automation-first thinking essential to success.

Skills: bgp,design,nvidia,rdma,networking,infiniband,python,ansible,ip,automation

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Engineer — AI Datacenter
Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Principal Network Engineer — AI Datacenter
Principal Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Principal Network Engineer — GPU Infrastructure
Principal Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Network Engineer — GPU Infrastructure
Senior Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Network Engineer — GPU Infrastructure
Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 900,000 - 1,300,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Gruppe • Pune District

On-site
INR 1,200,000 - 2,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA Gruppe • Gurugram District

On-site
INR 3,000,000 - 6,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000