Infrastructure SRE - HPC

Sarvam AI

Bengaluru

On-site

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sarvam AI in Bengaluru is hiring a specialist to manage its multi-vendor GPU fleet. The role demands operating GPU clusters end to end, maintaining system performance, and collaborating with teams to ensure reliability during high-demand workloads.

Ideal candidates have over 5 years of experience in infrastructure engineering, with specific expertise in GPU clusters. The position offers the chance to work on cutting-edge technology and contribute to building AI solutions in India.

Qualifications

  • 5+ years in infrastructure or site reliability engineering.
  • 2+ years operating GPU clusters at scale.
  • Demonstrated ownership of critical infrastructure.

Responsibilities

  • Operate the GPU fleet across training and serving.
  • Participate in on-call rotations and write runbooks.
  • Build internal tooling for team efficiency.
  • Collaborate with ML and platform teams.

Skills

Infrastructure or site reliability engineering
Operating GPU clusters at scale
Python
Go

Tools

Slurm
Kubernetes

Job description

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

Sarvam runs a large, multi-vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs and must run uninterrupted for weeks, and inference services that must hold a flat p99 under production load. Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve.

This is not a Kubernetes administration role. We assume Kubernetes fluency as a baseline. The difficulty lies above and below it - in parallel filesystems under heavy checkpoint load, in RDMA fabrics that degrade quietly, in NCCL hangs whose root cause may be the network or the kernel, in driver and firmware drift across heterogeneous hardware, and in distributed training failures that masquerade as infrastructure faults.

We are hiring a team of specialists rather than a set of identical generalists. This posting covers five areas of focus. We expect candidates to bring genuine depth in one and working fluency across the others, because on a shared fleet a storage problem often first appears as a training hang, and the engineer on call must route an incident correctly before anyone can resolve it.

What You’ll Do
  • Operate the GPU fleet end to end across training and serving - provisioning, observability, capacity, and fleet health.
  • Hold a meaningful on-call rotation, write runbooks that hold up under pressure, and drive postmortems that produce durable fixes.
  • Build the internal tooling the team relies on, rather than operating off-the-shelf systems alone.
  • Partner with ML and platform teams to keep large runs alive and serving latency predictable.
What We're Looking For
  • 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.
  • Demonstrated on-call ownership of infrastructure that mattered, with a track record of postmortems that led to real change.
  • Proficiency in Python or Go, used to build and maintain internal tooling.
  • Working fluency across all five areas of focus below - enough to recognize, triage, and route a problem outside your specialty, even if the fix belongs to a teammate.

* For the Storage and Fabric areas of focus, we will weigh deep domain expertise against the GPU-cluster requirement; exceptional specialists with less direct GPU-fleet time are encouraged to apply.

Bonus Points
  • Slurm and Kubernetes hybrid environments.
  • On-premise GPU deployment, including coordination with datacenter operations on power, cooling, and InfiniBand cabling.
  • Experience with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud etc.
  • Multi-tenant GPU isolation (MIG, MPS, time-slicing) in production.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure SRE - HPC
Infrastructure SRE - HPC

Sarvam • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Platform Engineer - AI Infrastructure
Platform Engineer - AI Infrastructure

Sarvam • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Health benefits
Flexible working hours
Platform Engineer - AI Infrastructure
Platform Engineer - AI Infrastructure

Sarvam • Chennai District

On-site
INR 4,000,000 - 7,000,000
Embedded Infrastructure Engineer, Chanakya
Embedded Infrastructure Engineer, Chanakya

Sarvam • Delhi

On-site
INR 1,400,000 - 2,000,000
ML Engineer (Training Infra), Foundational Models
ML Engineer (Training Infra), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,200,000 - 2,000,000
High ownership and impact
AI-first approach
Embedded Infrastructure Engineer, Chanakya
Embedded Infrastructure Engineer, Chanakya

Neara • Delhi

On-site
INR 1,200,000 - 2,000,000
Impactful work
High ownership
Collaborative team environment
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Performance Engineer, On-Device Inference
Performance Engineer, On-Device Inference

Sarvam • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Staff Engineer, Product Infrastructure
Staff Engineer, Product Infrastructure

Sarvam • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Performance Engineer, Intel Stack
Senior Performance Engineer, Intel Stack

Sarvam • Bengaluru

On-site
INR 1,500,000 - 2,500,000