GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro

Mumbai

On-site

INR 3,600,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Larsen & Toubro is seeking an expert to build and operate large-scale GPU compute pods for predictable training and inference services across a 10K GPU cluster. You will stand up multi-pod clusters, implement MIG/vGPU quotas, and integrate Slurm/Kubernetes with device plugins and accounting.

You will own 24/7 operations, capacity planning, and performance tuning to meet stringent SLAs. You will drive GPU lifecycle from firmware to driver, coordinate cross-functional teams, and lead RCA for

Qualifications

  • Experience in building and operating HPC/AI GPU clusters at scale.
  • Deep CUDA/MIG/vGPU expertise.
  • Proven with H100 GPUs and Slurm/K8s device scheduling; NCCL/UCX performance tuning.
  • Strong incident response and RCA capabilities.

Responsibilities

  • Stand up multi-pod GPU clusters with rack/power/cooling layouts and connectivity.
  • Implement MIG/vGPU quotas for multi-tenant environments.
  • Integrate cluster schedulers (Slurm/Kubernetes) with device plugins and accounting.
  • Own day-to-day operations across firmware/driver/DCGM/NVML lifecycles with minimal downtime.
  • Capacity planning for GPUs/CPU/Memory/NIC and bin-packing across heterogeneous SKUs.
  • Tune CUDA runtimes, GPU clocks, GPUDirect Storage, and NUMA locality for performance.
  • Lead P0/P1 incident responses and coordinate cross-functional teams to restore services within SLA.
  • Perform RCA for network, GPU, CUDA, scheduler failures and implement preventive actions.
  • Collaborate with platform/GPU teams to resolve CUDA job and scheduling issues in large-scale AI training.

Skills

CUDA tuning
NCCL/UCX
Slurm scheduling
Kubernetes scheduling
GPU virtualization
Cluster capacity planning
Incident management
Root cause analysis

Education

BE/B-Tech in CS/EC
NVIDIA Certified Professional (AI Infrastructure)
Linux (RHCE/LPIC)

Tools

NVIDIA GPUs (H100)
Slurm
Kubernetes
NVIDIA DCGM/NVML
nvidia-smi
GPUDirect Storage
Prometheus/Grafana
ELK/Splunk

Job description

Job Purpose

Build and operate large?scale GPU compute pods to deliver predictable, high?throughput, low?latency training and inference services across 10K GPU cluster.

Roles & Responsibilities
Implementation
  • Stand up multi?pod GPU clusters (rack/power/cooling layouts; TOR/leaf connectivity; IB/Ethernet host configs).
  • Implement GPU partitioning (MIG/vGPU profiles) and quota policies for multi?tenant environments.
  • Integrate cluster schedulers (Slurm/Kubernetes) with GPU device plugins, node feature discovery, and accounting/quotas.
Operations
  • Own day?2 operations across firmware/driver/DCGM/NVML lifecycles; execute change windows with zero/low downtime.
  • Capacity planning (GPU/CPU/Memory/NIC) and bin?packing strategies for heterogeneous GPU SKUs.
Performance & Optimization
  • Tune NCCL/UCX, GPU clocks/persistence, GPU Direct Storage, NUMA/locality, and CUDA runtime parameters.
  • Drive benchmarking and acceptance (HPL, HPL?AI, MLPerf?like internal suites); track perf regressions with SLOs.
Reliability & Incident
  • Lead P0/P1 incident response for GPU, CUDA, or scheduler issues; perform root?cause and preventative actions.
  • Act as the primary technical lead during production outages, coordinating cross-functional teams across Networking, GPU Operations, Platform Engineering, Storage, and Application teams to restore services within SLA targets.
  • Perform detailed Root Cause Analysis (RCA) for network, GPU, CUDA, NCCL, RDMA, and scheduler-related failures, identifying underlying causes and implementing preventive and corrective actions.
  • Collaborate with platform and GPU engineering teams to resolve issues impacting CUDA jobs, Kubernetes scheduling, Slurm workload management, GPU resource allocation, and large-scale AI training environments.
  • Define golden images; implement node remediation (cordon/drain/reimage) and auto?healing workflows.
Security & Compliance
  • Enforce GPU tenancy isolation (MIG, cgroup/device cgroup, mpsd), secure drivers/containers, SBOM and image scanning.
Documentation & Enablement
  • Publish runbooks, performance baselines, and application tuning guides for LLM training and inference.
Experience & Educational Requirement
  • BE/B-Tech or equivalent with Computer Science or Electronics & Communication
  • NVIDIA Certified Professional (AI Infrastructure); Linux (RHCE/LPIC) preferred.
Relevant Experience
  • 7–12 years building/operating HPC/AI GPU clusters at scale; deep CUDA/MIG/vGPU expertise.
  • Proven with H100/B?series class GPUs, Slurm and/or Kubernetes device scheduling; NCCL/UCX performance tuning.
Tools / Tech
  • CUDA, NCCL, cuDNN, TensorRT; 2. DCGM/NVML, nvidia-smi; 3. Slurm, K8s (device plugin, DaemonSets, NFD); 4. Helm/ArgoCD; 6. GPUDirect Storage; 7. Prometheus/Grafana; 8. ELK/Splunk.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
GPU Cluster Architect
GPU Cluster Architect

Nebius B.V. • India

On-site
INR 4,000,000 - 6,500,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead Engineer (HPC, GPU, CUDA)
Lead Engineer (HPC, GPU, CUDA)

AIRA Matrix • Thane

On-site
INR 1,800,000 - 3,200,000
AI Engineer
AI Engineer

STCO India • Hyderabad

On-site
INR 800,000 - 1,200,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA AI • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA • Bengaluru

On-site
INR 2,500,000 - 4,000,000