Principal Network Engineer — AI Datacenter

Nava

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Nava designs, deploys and operates GPU infrastructure and offers inference-as-a-service for AI teams. The Network Engineering team owns the datacenter fabric lifecycle, from architecture to production operation, focusing on scalable, low-latency, lossless networks for distributed training and inference.

We seek an experienced architect/engineer to design RoCEv2 and EVPN-VXLAN leaf-spine fabrics, drive automation with Python, Ansible and Terraform, and mentor senior engineers while remaining

Qualifications

  • 8–10 years designing and operating datacenter networks at scale.
  • Hands-on command of BGP, OSPF, IS-IS, TCP/IP, IPv4/IPv6, DNS, DHCP and MPLS.
  • Experience with RDMA fabrics (RoCEv2/InfiniBand) and EVPN-VXLAN.
  • Automation skills in Python plus Ansible and/or Terraform.
  • Comfortable operating in a fast-paced, on-call production environment.

Responsibilities

  • Own architecture and high-level design of Nava's GPU cluster fabrics (RoCEv2, leaf-spine, datacenter interconnect).
  • Define fleet-wide standards for BGP/EVPN-VXLAN topologies and congestion control.
  • Drive network automation and decompose designs into buildable implementations.
  • Lead cross-functional initiatives with Compute, Storage, SRE, and security teams.
  • Mentor engineers and participate in hands-on cluster turn-ups and incident response.

Skills

BGP
OSPF
IS-IS
TCP/IP
IPv6
DNS
DHCP
MPLS
RDMA
RoCEv2
EVPN-VXLAN
Python
Linux

Tools

Ansible
Terraform
Cisco
Arista
NVIDIA Spectrum-X

Job description

About Nava

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.

About Nava

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.

About The Team

The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, from architecture and design through deployment, automation, and production operation. We build the lossless, high-throughput networks that carry GPU-to-GPU traffic for distributed training and low-latency inference. This is a hands-on team where the network is treated as code—and where your expertise will shape the foundation of AI infrastructure at scale.

Responsibilities
  • Own the architecture and high-level design of Nava's GPU cluster fabrics, including RoCEv2 backend networks, frontend networks, and datacenter interconnect.
  • Define fleet-wide standards for BGP and EVPN-VXLAN leaf-spine topologies, congestion control (PFC/ECN tuning), and lossless RDMA transport.
  • Set the technical direction for network automation, decomposing high-level architecture into detailed, buildable designs.
  • Lead cross-functional initiatives with Compute, Storage, SRE, and security teams to deliver fabrics that never bottleneck training or inference.
  • Drive root-cause analysis on the hardest fabric performance and reliability problems and establish preventive practices.
  • Mentor Senior and Network Engineers, review designs, and raise the technical bar across the team.
  • Remain hands-on: lead cluster turn-ups, acceptance testing, RoCEv2 implementation, incident response, and PFC/ECN tuning under real load, alongside architecture and strategy.
Required Qualifications
  • 8–10 years designing and operating datacenter networks at scale, with protocol-level architecture authority and a track record leading complex, multi-vendor designs.
  • Strong hands-on command of core networking protocols: BGP, OSPF, IS-IS, TCP/IP, IPv4 and IPv6, DNS, DHCP, and MPLS.
  • Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL/TLS.
  • Solid understanding of datacenter fabric concepts (leaf-spine, EVPN-VXLAN) and RDMA fabrics (RoCEv2 and InfiniBand), including lossless Ethernet.
  • Automation skills: Python plus Ansible and/or Terraform.
  • Hands-on experience with RDMA fabrics (RoCEv2 and/or InfiniBand), including congestion management for lossless Ethernet and IB fabrics.
  • Comfortable operating in a fast-paced, on-call production environment.
Preferred Qualifications
  • Experience with NVIDIA networking (Spectrum-X, Quantum InfiniBand, BlueField DPUs) and NCCL traffic patterns.
  • Experience operating GPU clusters for large-scale distributed training or inference.
  • Familiarity with network telemetry, streaming analytics, and closed-loop automation.
  • Relevant certifications (e.g., CCNP or vendor equivalents).
  • Experience with Cisco and/or Arista platforms is a strong plus.
Technology environment

NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, NVIDIA networking (Spectrum/Quantum, BlueField DPUs); RoCEv2 and InfiniBand RDMA lossless fabrics; BGP, OSPF, IS-IS, EVPN-VXLAN, MPLS; TCP/IP, IPv4/IPv6, DNS, DHCP, VPN, SSL/TLS; Ansible, Terraform, Python; Linux at scale.

Skills: automation,nvidia,design,infiniband,ip,training,tcp/ip,rdma,networking,architecture

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Engineer — AI Datacenter
Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Network Engineer — AI Datacenter
Senior Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Senior Network Engineer — GPU Infrastructure
Senior Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Principal Network Engineer — GPU Infrastructure
Principal Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Network Engineer — GPU Infrastructure
Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 900,000 - 1,300,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Senior Network Engineer
Senior Network Engineer

Operations • Mumbai

On-site
INR 600,000 - 1,200,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Gruppe • Pune District

On-site
INR 1,200,000 - 2,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000
Senior Network Infrastructure Engineer
Senior Network Infrastructure Engineer

NVIDIA • Hyderabad

On-site
INR 1,800,000 - 2,700,000