Principal Network Engineer — GPU Infrastructure

Nava

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nava designs, deploys, and operates large-scale GPU infrastructure and delivers inference-as-a-service for AI products. The Compute team owns GPU systems, software, networking, and day-to-day operations to keep thousands of GPUs healthy and running at peak performance.

We seek an experienced infrastructure engineer to architect, implement, and automate tasks across GPU clusters, leveraging RDMA, NVLink, and the NVIDIA stack for reliable, scalable AI workloads.

Qualifications

  • 8–10 years of infra engineering with leadership in large-scale compute systems.
  • Hands-on Kubernetes administration and architectural experience.
  • Experience deploying GPU clusters in production environments.
  • Strong Linux knowledge and NVIDIA GPU stack including CUDA and RDMA.
  • Experience with Ansible and Terraform for automation.
  • Familiarity with monitoring tools like Prometheus, Grafana, Datadog, Dynatrace, or Zabbix.

Responsibilities

  • Design and evolve architecture for large-scale NVIDIA GPU clusters connected via NVLink and RoCEv2/InfiniBand.
  • Define standards for bare-metal provisioning, GPU node setup, and cluster lifecycle automation.
  • Lead automation across the fleet using Python, Ansible, and Terraform to deploy, configure, and maintain GPU clusters at scale.
  • Build and maintain GPU health monitoring, fault detection, and fast recovery when nodes fail.
  • Collaborate with Network and Storage teams to ensure data paths are optimized and GPUs stay fully utilized.
  • Mentor senior and junior engineers and help set technical standards for the team.

Skills

Kubernetes administration
Linux systems
GPU infrastructure
Infrastructure automation
Monitoring & observability
RDMA/NVLink/NVIDIA stack
Leadership & mentoring
Cluster architecture

Tools

Ansible
Terraform
Prometheus
Grafana
NVIDIA GPU Operator
Kubectl
Slurm
NCCL

Job description

About Nava

Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.


About the team

The Compute team owns Nava's GPU infrastructure end-to-end from installing and configuring NVIDIA GPU systems, to managing cluster software, networking, and day-to-day operations. Our goal: keep thousands of GPUs healthy, responsive, and running at peak performance for training and inference workloads.


Responsibilities


  • Design and evolve the architecture for large-scale NVIDIA GPU clusters connected via NVLink and high-speed networking (RoCEv2/InfiniBand).

  • Define standards for bare-metal provisioning, GPU node setup (drivers, CUDA, GPU Operator, RDMA), and cluster lifecycle automation.

  • Lead automation efforts across the fleet using Python, Ansible, and Terraform to deploy, configure, and maintain GPU clusters at scale.

  • Build and maintain practices for GPU health monitoring, fault detection, and fast recovery when nodes fail.

  • Collaborate with Network and Storage teams to ensure data paths are optimized and GPUs stay fully utilized.

  • Mentor both senior and junior engineers and help set technical standards for the team.


Qualifications


  • Hands-on experience with Kubernetes administration and architecture.

  • Experience deploying and managing GPU clusters in production environments.

  • Strong understanding of Kubernetes networking, storage, and security.

  • Experience building infrastructure automation using tools like Ansible and Terraform.

  • Deep Linux systems knowledge and familiarity with the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).

  • 8-10 years in infrastructure engineering with recent experience in compute infrastructure and a history of leading architecture for large-scale systems.

  • Experience with NVIDIA GPU systems (e.g., DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator, and RDMA fabrics.

  • Experience with monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.

  • Familiarity with Slurm, NCCL, NVLink tuning, and distributed training/inference workloads is a plus.


Preferred qualifications


  • Experience in AI/ML infrastructure.

  • Experience with hybrid cloud and datacenter infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network Engineer — GPU Infrastructure
Senior Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Network Engineer — GPU Infrastructure
Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 900,000 - 1,300,000
Senior Network Engineer — AI Datacenter
Senior Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Principal Network Engineer — AI Datacenter
Principal Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Network Engineer — AI Datacenter
Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA Gruppe • Gurugram District

On-site
INR 3,000,000 - 6,000,000