Get more replies from employers
Send a job-specific resume in minutes.
Sarvam is seeking an accomplished infrastructure/SRE engineer to govern a large GPU fleet used for training and inference. You will manage provisioning, observability, capacity, and fleet health while owning on-call incidents and crafting durable runbooks.
You’ll build internal tooling and collaborate with ML and platform teams to keep workloads running with predictable latency. Ideal candidates bring 5+ years in infra or SRE, including GPU cluster experience, Python or Go proficiency, and a
Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.
Sarvam runs a large, multi-vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs and must run uninterrupted for weeks, and inference services that must hold a flat p99 under production load. Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve.
This is not a Kubernetes administration role. We assume Kubernetes fluency as a baseline. The difficulty lies above and below it - in parallel filesystems under heavy checkpoint load, in RDMA fabrics that degrade quietly, in NCCL hangs whose root cause may be the network or the kernel, in driver and firmware drift across heterogeneous hardware, and in distributed training failures that masquerade as infrastructure faults.
We are hiring a team of specialists rather than a set of identical generalists. This posting covers five areas of focus. We expect candidates to bring genuine depth in one and working fluency across the others, because on a shared fleet a storage problem often first appears as a training hang, and the engineer on call must route an incident correctly before anyone can resolve it.
Bring depth in one of the five areas below; expect to be conversational across the rest.
Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.
Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar.
High ownership and high impact, from day one.
Everything we do is AI-first, from the way we build and ship to the way we think about problems.
You can work on problems that could change how an entire country learns, works, and communicates.
If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.