Get more replies from employers
Send a job-specific resume in minutes.
Larsen & Toubro is seeking an expert to build and operate large-scale GPU compute pods for predictable training and inference services across a 10K GPU cluster. You will stand up multi-pod clusters, implement MIG/vGPU quotas, and integrate Slurm/Kubernetes with device plugins and accounting.
You will own 24/7 operations, capacity planning, and performance tuning to meet stringent SLAs. You will drive GPU lifecycle from firmware to driver, coordinate cross-functional teams, and lead RCA for
Build and operate large?scale GPU compute pods to deliver predictable, high?throughput, low?latency training and inference services across 10K GPU cluster.