Senior Network Engineer — GPU Infrastructure

Nava

Bengaluru

On-site

INR 3,000,000 - 6,000,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Nava is building out its Compute team to design, deploy, and operate large-scale GPU infrastructure powering AI inference at scale. The role centers on hands-on GPU cluster deployment, Kubernetes-centric operations, and end-to-end performance tuning in a production environment.

You will drive automation, contribute to GitOps pipelines, and collaborate with Network and Storage teams while mentoring junior engineers.

Qualifications

  • 5–8 years in compute/systems infrastructure with hands-on GPU cluster experience.
  • Deep knowledge of Kubernetes networking, storage, and security.
  • Experience with infrastructure automation tools (Ansible, Terraform).
  • Strong Linux systems expertise and NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA.

Responsibilities

  • Deploy and operate NVIDIA GPU clusters: bare-metal provisioning, node configuration, driver/CUDA/RDMA setup, and validation.
  • Build automation in Python, Ansible, and Terraform to provision and maintain clusters and reduce manual toil.
  • Extend Kubernetes using custom controllers and operators—design and develop custom Kubernetes controllers and operators.
  • Build and maintain GitOps-based deployments, CI/CD pipelines, deployment automation, and infrastructure tooling.
  • Integrate GPU nodes with RoCEv2/InfiniBand fabrics and validate end-to-end performance (NCCL, benchmarks).
  • Lead new cluster turn-ups and acceptance testing; troubleshoot GPU, node, and interconnect issues.
  • Improve monitoring and automated recovery to keep utilization high and MTTR low.
  • Collaborate with Network and Storage teams and mentor junior engineers.

Skills

Kubernetes
GPU clusters
Linux
CUDA stack
RDMA networking
Ansible
Terraform
NVIDIA software stack

Tools

Ansible
Terraform

Job description

About Nava

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.

About The Team

The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.

Responsibilities
  • Deploy and operate NVIDIA GPU clusters: bare-metal provisioning, node configuration, driver/CUDA/RDMA setup, and validation.
  • Build automation in Python, Ansible, and Terraform to provision and maintain clusters and reduce manual toil.
  • Extend Kubernetes using custom controllers and operators—design and develop custom Kubernetes controllers and operators.
  • Build and maintain GitOps-based deployments, CI/CD pipelines, deployment automation, and infrastructure tooling.
  • Integrate GPU nodes with RoCEv2/InfiniBand fabrics and validate end-to-end performance (NCCL, benchmarks).
  • Lead new cluster turn-ups and acceptance testing; troubleshoot GPU, node, and interconnect issues.
  • Improve monitoring and automated recovery to keep utilization high and MTTR low.
  • Collaborate with Network and Storage teams and mentor junior engineers.
Required Qualifications
  • Strong hands-on experience with Kubernetes administration and architecture.
  • GPU cluster deployment and management experience.
  • Deep understanding of Kubernetes networking, storage, and security.
  • Experience with infrastructure automation tools (e.g., Ansible, Terraform).
  • Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
  • 5–8 years in compute/systems infrastructure with hands-on GPU cluster experience; comfortable in an on-call production environment.
Preferred Qualifications
  • Exposure to the NVIDIA GPU ecosystem.
  • Experience with AI/ML infrastructure.
  • Familiarity with monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
  • Experience with hybrid cloud and datacenter infrastructure.
  • Hands-on experience with NCCL and NVLink tuning, Slurm, and distributed training/inference workloads.
Technology environment

NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator; RoCEv2 / InfiniBand RDMA fabrics; Kubernetes (networking, storage, security), Slurm; Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix; Linux at scale.

Skills

metal,kubernetes,automation,nvidia,infiniband,cuda,cluster,infrastructure,rdma,ansible

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Engineer — GPU Infrastructure
Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 900,000 - 1,300,000
Principal Network Engineer — GPU Infrastructure
Principal Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Principal Network Engineer — AI Datacenter
Principal Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Network Engineer — AI Datacenter
Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Network Engineer — AI Datacenter
Senior Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000