Principal Network Engineer — GPU Infrastructure

Nava

Bengaluru

On-site

INR 4,000,000 - 6,500,000

Full time

9 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nava is building and operating large-scale GPU infrastructure in Bengaluru to power AI workloads. The Compute team handles NVIDIA GPU systems, cluster software, networking, and day-to-day operations to keep thousands of GPUs healthy and responsive for training and inference workloads.

The role requires deep Linux knowledge, NVIDIA stack experience, and leading architecture for compute infrastructure. Hybrid cloud and data center exposure are a plus, with mentoring responsibilities across the

Qualifications

  • Hands-on Kubernetes administration and architecture.
  • Experience deploying and managing GPU clusters in production.
  • Strong understanding of Kubernetes networking, storage, and security.
  • Experience building infrastructure automation using Ansible and Terraform.
  • Deep Linux systems knowledge with NVIDIA stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
  • 8–10 years in infrastructure engineering with leadership of large-scale systems architecture.

Responsibilities

  • Design and evolve the architecture for large-scale NVIDIA GPU clusters connected via NVLink and high-speed networking (RoCEv2/InfiniBand).
  • Define standards for bare-metal provisioning, GPU node setup, and cluster lifecycle automation.
  • Lead automation across the fleet using Python, Ansible, and Terraform to deploy, configure, and maintain GPU clusters at scale.
  • Build and maintain practices for GPU health monitoring, fault detection, and fast recovery when nodes fail.
  • Collaborate with Network and Storage teams to ensure data paths are optimized and GPUs stay fully utilized.
  • Mentor both senior and junior engineers—and help set technical standards for the team.

Skills

Kubernetes admin
GPU clusters
Automation tooling
RDMA/InfiniBand

Tools

Ansible
Terraform

Job description

About Nava

Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure-and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.

About Nava

Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure-and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.

About The Team

The Compute team owns Nava's GPU infrastructure end-to-end-from installing and configuring NVIDIA GPU systems, to managing cluster software, networking, and day-to-day operations. Our goal: keep thousands of GPUs healthy, responsive, and running at peak performance for training and inference workloads.

Responsibilities
  • Design and evolve the architecture for large-scale NVIDIA GPU clusters connected via NVLink and high-speed networking (RoCEv2/InfiniBand).
  • Define standards for bare-metal provisioning, GPU node setup (drivers, CUDA, GPU Operator, RDMA), and cluster lifecycle automation.
  • Lead automation efforts across the fleet using Python, Ansible, and Terraform to deploy, configure, and maintain GPU clusters at scale.
  • Build and maintain practices for GPU health monitoring, fault detection, and fast recovery when nodes fail.
  • Collaborate with Network and Storage teams to ensure data paths are optimized and GPUs stay fully utilized.
  • Mentor both senior and junior engineers-and help set technical standards for the team.
Qualifications
  • Hands-on experience with Kubernetes administration and architecture.
  • Experience deploying and managing GPU clusters in production environments.
  • Strong understanding of Kubernetes networking, storage, and security.
  • Experience building infrastructure automation using tools like Ansible and Terraform.
  • Deep Linux systems knowledge-and familiarity with the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
  • 8-10 years in infrastructure engineering-with recent experience in compute infrastructure and a history of leading architecture for large-scale systems.
  • Experience with NVIDIA GPU systems (e.g., DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator, and RDMA fabrics.
  • Experience with monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
  • Familiarity with Slurm, NCCL, NVLink tuning, and distributed training/inference workloads is a plus.
Preferred Qualifications
  • Experience in AI/ML infrastructure.
  • Experience with hybrid cloud and datacenter infrastructure.

Skills: nvidia,kubernetes,cluster,rdma,infiniband,cuda,architecture,ansible,infrastructure,automation

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of GPU Cluster Engineering
Head of GPU Cluster Engineering

Nava • Bengaluru

On-site
INR 6,000,000 - 11,000,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior System Software Engineer, Software Defined Networking
Senior System Software Engineer, Software Defined Networking

NVIDIA Corporation • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000
NCX Senior Engineer
NCX Senior Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Senior High-Performance Storage Architect - NVIS
Senior High-Performance Storage Architect - NVIS

NVIDIA Corporation • India

On-site
INR 14,231,000 - 19,924,000