Network Engineer — GPU Infrastructure

Nava

Bengaluru

On-site

INR 900,000 - 1,300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nava in Bengaluru is seeking a capable Systems Engineer to design, deploy, and operate GPU infrastructure for AI workloads. You will manage bare‑metal GPU clusters and ensure fast, reliable training and inference at scale.

You will implement OS provisioning, CUDA drivers, GPU Operator, and RDMA networking; write automation with Python, Ansible, and Terraform; run health checks, benchmarks, and documentation; collaborate with cross‑functional teams to meet technical requirements.

Qualifications

  • Strong hands-on Kubernetes administration and architecture.
  • Bare Metal GPU cluster deployment and management experience.
  • Kubernetes networking, storage, and security.
  • Infrastructure automation (Ansible, Terraform, etc.).
  • Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers) and RDMA.
  • 2-4 years in systems/infrastructure engineering with Linux fundamentals.

Responsibilities

  • Provision and configure GPU nodes: OS, drivers, CUDA, GPU Operator, and RDMA connectivity.
  • Write and maintain automation using Python, Ansible, and Terraform.
  • Support cluster deployments, run health checks and benchmarks, and document results.
  • Monitor GPU infrastructure, respond to alerts, and resolve or elevate issues.
  • Maintain runbooks and operational documentation.
  • Collaborate with cross-functional engineering teams to deliver software solutions.

Skills

Kubernetes
Bare metal GPU
Networking
Storage
Security
Automation
Linux
CUDA
NVIDIA drivers
2-4 years experience

Tools

Ansible
Terraform

Job description

About Nava

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the bare metal, on infrastructure built to keep GPUs available and models serving.

About the team

The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.

Responsibilities
  • Provision and configure GPU nodes: OS, drivers, CUDA, GPU Operator, and RDMA connectivity.
  • Write and maintain automation using Python, Ansible, and Terraform.
  • Support cluster deployments, run health checks and benchmarks, and document results.
  • Monitor GPU infrastructure, respond to alerts, and resolve or elevate issues.
  • Maintain runbooks and operational documentation.
  • Collaborate with cross-functional engineering teams to understand business and technical requirements and deliver software solutions.
Required qualifications
  • Strong hands-on experience with Kubernetes administration and architecture.
  • Bare Metal GPU cluster deployment and management experience.
  • Kubernetes networking, storage, and security.
  • Infrastructure automation (Ansible, Terraform, etc.).
  • Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
  • 2-4 years in systems/infrastructure engineering with strong Linux fundamentals and a learning-oriented mindset.
Preferred qualifications
  • NVIDIA GPU ecosystem exposure.
  • AI/ML infrastructure experience.
  • Monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
  • Hybrid cloud and datacenter infrastructure experience.
  • NCCL and NVLink tuning, Slurm, and distributed training/inference workloads.
Technology environment

NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator; RoCEv2 / InfiniBand RDMA fabrics; Kubernetes (networking, storage, security), Slurm; Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix; Linux at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network Engineer — GPU Infrastructure
Senior Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Principal Network Engineer — GPU Infrastructure
Principal Network Engineer — GPU Infrastructure

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Network Engineer — AI Datacenter
Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Principal Network Engineer — AI Datacenter
Principal Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Network Engineer — AI Datacenter
Senior Network Engineer — AI Datacenter

Nava • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000