GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro-Vyoma

Chennai District

On-site

INR 3,000,000 - 6,000,000

Full time

10 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Runbooks
Performance baselines
Application tuning guides

Job summary

Larsen & Toubro-Vyoma is seeking an experienced professional to build and operate large-scale GPU compute pods for high-throughput AI training and inference services in Chennai. You will design rack/power/cooling layouts, implement MIG/vGPU quotas, and integrate Slurm/Kubernetes with device plugins and accounting.

You will own day-2 operations, perform capacity planning for heterogeneous GPU SKUs, and drive performance tuning of NCCL/UCX, CUDA routines, and GPU clocks.

Qualifications

  • BE/B-Tech or equivalent with Computer Science or Electronics & Communication
  • 7-12 years building/operating HPC/AI GPU clusters at scale; CUDA/MIG/vGPU expertise
  • Proven with H100/B-series GPUs, Slurm and/or Kubernetes device scheduling; NCCL/UCX performance tuning.

Responsibilities

  • Stand up multi-pod GPU clusters with GPU partitioning policies.
  • Own day-2 operations across firmware/driver/DCGM/NVML lifecycles; zero/low downtime.
  • Lead P0/P1 incident response for GPU, CUDA, or scheduler issues.
  • Define golden images; implement node remediation (cordon/drain/reimage) and auto-healing workflows.

Skills

CUDA programming
MIG/vGPU
Slurm scheduling
Kubernetes
NCCL/UCX
GPU hardware

Education

BE/BTech in CS or ECE

Tools

Slurm
Kubernetes
Docker

Job description

Build and operate large-scale GPU compute pods to deliver predictable, high-throughput, low-latency training and inference services across 10K GPU cluster.

Implementation

Stand up multi-pod GPU clusters (rack/power/cooling layouts; TOR/leaf connectivity; IB/Ethernet host configs).

Implement GPU partitioning (MIG/vGPU profiles) and quota policies for multi-tenant environments.

Integrate cluster schedulers (Slurm/Kubernetes) with GPU device plugins, node feature discovery, and accounting/quotas.

Operations

Own day-2 operations across firmware/driver/DCGM/NVML lifecycles; execute change windows with zero/low downtime.

Capacity planning (GPU/CPU/Memory/NIC) and bin-packing strategies for heterogeneous GPU SKUs.

Performance & Optimization

Tune NCCL/UCX, GPU clocks/persistence, GPU Direct Storage, NUMA/locality, and CUDA runtime parameters.

Drive benchmarking and acceptance (HPL, HPL-AI, MLPerf-like internal suites); track perf regressions with SLOs.

Reliability & Incident

Lead P0/P1 incident response for GPU, CUDA, or scheduler issues; perform root-cause and preventative actions.

Act as the primary technical lead during production outages, coordinating cross-functional teams across Networking, GPU Operations, Platform Engineering, Storage, and Application teams to restore services within SLA targets.

Perform detailed Root Cause Analysis (RCA) for network, GPU, CUDA, NCCL, RDMA, and scheduler-related failures, identifying underlying causes and implementing preventive and corrective actions.

Collaborate with platform and GPU engineering teams to resolve issues impacting CUDA jobs, Kubernetes scheduling, Slurm workload management, GPU resource allocation, and large-scale AI training environments.

Define golden images; implement node remediation (cordon/drain/reimage) and auto-healing workflows.

Security & Compliance

Enforce GPU tenancy isolation (MIG, cgroup/device cgroup, mpsd), secure drivers/containers, SBOM and image scanning.

Documentation & Enablement

Publish runbooks, performance baselines, and application tuning guides for LLM training and inference.

Experience & Educational Requirement

BE/B-Tech or equivalent with Computer Science or Electronics & Communication

RELEVANT EXPERIENCE
  • 7-12 years building/operating HPC/AI GPU clusters at scale; deep CUDA/MIG/vGPU expertise.
  • Proven with H100/B-series class GPUs, Slurm and/or Kubernetes device scheduling; NCCL/UCX performance tuning.
Tools / Tech
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Platform Engineer
GPU Platform Engineer

Lancesoft • Dadri

On-site
INR 3,500,000 - 6,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
GPU Cluster Architect
GPU Cluster Architect

Nebius B.V. • India

On-site
INR 4,000,000 - 6,500,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
GPU Infrastructure Engineer
GPU Infrastructure Engineer

TECEZE • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Principal Software Architect- High Performance Computing
Principal Software Architect- High Performance Computing

Applied Materials India • Chennai District

On-site
INR 4,000,000 - 6,000,000
AI Compute Engineer
AI Compute Engineer

DAMAC Digital • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA • Bengaluru

On-site
INR 2,500,000 - 4,000,000