GPU Platform Engineer

Lancesoft

Dadri

On-site

INR 3,500,000 - 6,000,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

LanceSoft is expanding its AI infrastructure practice and seeks a GPU Platform Engineer to design, deploy and operate GPU-accelerated platforms for enterprise clients. You will own health, performance and scalability from deployment through steady-state operations.

Work includes NVIDIA DGX/HGX clusters, on-prem and air-gapped deployments, multi-tenant pools, and integration with Slurm, Kubernetes, Docker and NGC containers. Strong Linux, scripting, and observability skills are essential.

Qualifications

  • Bachelor's degree in Computer Science, IT, Electronics or related field.
  • NCP-AIO certification current and valid.
  • Willingness to travel to client sites and join on-call rotations.

Responsibilities

  • Design and deploy GPU clusters on NVIDIA DGX/HGX platforms and OEM systems.
  • Configure Slurm and Kubernetes with NVIDIA GPU Operator and MIG.
  • Set up container workflows using Docker and NVIDIA NGC containers.
  • Validate InfiniBand/RoCE fabrics and NVLink/NVSwitch topologies.
  • Benchmark, tune and build observability with DCGM, Prometheus, Grafana.
  • Mentor NOC engineers and assist with pre-sales input.

Skills

NVIDIA AI stack
Linux administration
Slurm
Kubernetes
InfiniBand networking
Scripting (Python, Bash)
Observability tools
Communication

Education

Bachelor's degree in CS/IT/Electronics or related

Tools

NVIDIA Base Command Manager
NVIDIA NGC
Docker
Terraform
DCGM
NCCL

Job description

About the Role

LanceSoft is expanding its AI infrastructure practice and is looking for GPU Platform Engineers to design, build, operate and optimize GPU-accelerated platforms for enterprise clients. You will work on NVIDIA DGX and HGX based clusters, high-speed fabrics and AI workload platforms, and you will be the technical owner of platform health, performance and scalability from deployment through steady-state operations.

Required Certification
  • NVIDIA-Certified Professional: AI Operations (NCP-AIO). Must be current and valid at the time of joining.
Preferred certifications
  • NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • Cloud certifications (AWS, Azure or GCP)
  • CKA or RHCSA
  • Terraform Associate
Key Responsibilities-
Platform design and deployment
  • Design and deploy GPU clusters on NVIDIA DGX, HGX and OEM platforms (Dell, HPE, Supermicro and others), including compute, fabric and storage layers.
  • Provision and manage clusters using NVIDIA Base Command Manager or equivalent tooling.
  • Implement multi-tenant GPU platforms with workload isolation, quotas and fair-share scheduling.
  • Support both on-prem and air-gapped deployments, including offline software repositories and container registries.
Orchestration and workload management
  • Configure and operate Slurm and Kubernetes (including the NVIDIA GPU Operator, Network Operator and MIG configurations).
  • Set up and maintain container workflows using Docker, containerd and NVIDIA NGC containers.
  • Support AI and ML teams in onboarding training and inference workloads, and in right-sizing GPU allocation.
Networking and storage
  • Configure and validate InfiniBand and RoCE fabrics, NVLink and NVSwitch topologies, and manage fabrics through UFM.
  • Validate collective communication performance using NCCL tests and resolve bottlenecks.
  • Integrate high-throughput storage for AI (parallel file systems, NFS, object storage) and tune data pipelines.
Performance, reliability and observability
  • Benchmark and tune GPU, network and storage performance for training and inference workloads.
  • Build observability with DCGM, Prometheus and Grafana, and define alerts, dashboards and SLOs.
  • Diagnose GPU faults, XID errors, ECC and memory issues, thermal and power problems, and drive root cause analysis.
  • Act as the escalation point for the NOC and manage vendor cases with NVIDIA and OEM support.
Automation and lifecycle management
  • Automate provisioning, configuration and operations using Ansible, Terraform, Python and Bash.
  • Own firmware, driver, CUDA and software stack compatibility and upgrade planning.
  • Maintain reference architectures, runbooks, standards and platform documentation.
  • Plan capacity and forecast GPU demand with clients and internal teams.
Client and team engagement
  • Contribute technical input to client solution reviews, proposals and pre-sales discussions.
  • Mentor L1 to L3 NOC engineers and share operational best practice.
Required Skills
  • Deep understanding of the NVIDIA AI stack: DGX, Base Command, DCGM, NCCL, GPU Operator, NGC and NVIDIA AI Enterprise.
  • Strong Linux administration, including kernel, driver and package management.
  • Hands-on experience with Slurm and Kubernetes in GPU environments.
  • Working knowledge of InfiniBand, RoCE and high-performance networking.
  • Proficiency in scripting and infrastructure as code (Python, Bash, Ansible, Terraform).
  • Solid grasp of monitoring and observability tooling and incident troubleshooting.
  • Strong written and verbal communication, with the ability to explain platform decisions to technical and business stakeholders.
Qualifications
  • Bachelor's degree in Computer Science, IT, Electronics or a related field (or equivalent experience).
  • NCP-AIO certification, current and valid.
  • Willingness to travel to client sites and to join on-call rotations for critical escalations where required.
Nice to Have
  • Experience with NVIDIA Triton Inference Server, NeMo or other LLM serving and training stacks.
  • Experience with MLOps platforms and GPU cost or utilization management.
  • Experience with air-gapped or public sector AI environments.
  • Familiarity with liquid cooling, power and data center design considerations for high-density GPU racks.
  • ITIL 4 Foundation
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

NOC Engineer, AI Infrastructure (L1 / L2 / L3)
NOC Engineer, AI Infrastructure (L1 / L2 / L3)

Lancesoft • Dadri

On-site
INR 600,000 - 900,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
AI Engineer
AI Engineer

STCO India • Hyderabad

On-site
INR 800,000 - 1,200,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro-Vyoma • Chennai District

On-site
INR 3,000,000 - 6,000,000
Runbooks
Performance baselines
Application tuning guides
Specialist - System Management
Specialist - System Management

LTM • Chennai District

On-site
INR 1,200,000 - 1,800,000
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Corporation • Mumbai

On-site
INR 3,500,000 - 5,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000