AI Compute Engineer

DAMAC Digital

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

DAMAC Digital in Bengaluru is seeking an experienced HPC engineer to design and optimize GPU compute clusters for training and inference workloads. You will own architecture, deployment, and performance across bare metal, firmware, CUDA, NCCL, cuDNN, and the orchestration layer.

You'll integrate InfiniBand/RoCE fabrics, Slurm, Kubernetes and Run:AI, benchmark with NCCL tests and MLPerf, and lead vendor engagements with NVIDIA and partners.

Qualifications

  • 7+ years in HPC, AI or GPU infrastructure.
  • 3+ years hands-on with NVIDIA GPU clusters.
  • Deep NVIDIA stack expertise: CUDA, NCCL, cuDNN, TensorRT, DCGM.
  • InfiniBand/RoCE networking knowledge and Slurm/Kubernetes/Run:ai experience.
  • Advanced Linux, Python/Bash, Ansible or Terraform.

Responsibilities

  • Architect GPU compute solutions for large-scale training and inference clusters.
  • Deploy and tune NVIDIA Base Command Manager, CUDA, NCCL, cuDNN, TensorRT, DCGM, MIG.
  • Integrate with InfiniBand/RoCE fabrics and onboard Slurm, Kubernetes, Run:ai.
  • Benchmark with NCCL tests, MLPerf and HPL, and chase down bottlenecks.
  • Own firmware/driver upgrade cycles and lead vendor engagements with NVIDIA, Supermicro, Dell, HPE.

Skills

NVIDIA GPU clusters
CUDA / NCCL / cuDNN
Networking InfiniBand
Slurm / Kubernetes
Run:ai
Python / Bash
Ansible
Terraform
Advanced Linux
Vendor management

Tools

NVIDIA Base Command Manager
DCGM
MIG
MLPerf benchmarking

Job description

Someone has to design and tune the GPU clusters that actually run DAMAC AI's training and inference workloads. That's this role.

We're building one of the region's most ambitious AI compute footprints — NVIDIA B200, B300 and GB300 NVL72 clusters powering sovereign cloud and hyperscale AI services across the Middle East and Asia. You'll own architecture, deployment and performance across the full stack: bare metal, firmware, CUDA, NCCL, orchestration and multi-tenant scheduling.

What you'll do

  • Architect GPU compute solutions for large-scale training and inference clusters
  • Deploy and tune NVIDIA Base Command Manager, CUDA, NCCL, cuDNN, TensorRT, DCGM, MIG
  • Integrate with InfiniBand/RoCE fabrics and onboard Slurm, Kubernetes, Run:ai
  • Benchmark with NCCL tests, MLPerf and HPL, and chase down every bottleneck
  • Own firmware/driver upgrade cycles and lead vendor engagements with NVIDIA, Supermicro, Dell, HPE

What you bring

  • 7+ years in HPC, AI or GPU infrastructure, 3+ years hands-on with NVIDIA GPU clusters
  • Deep NVIDIA stack expertise: CUDA, NCCL, cuDNN, TensorRT, DCGM
  • InfiniBand/RoCE networking knowledge and Slurm/Kubernetes/Run:ai experience
  • Advanced Linux, Python/Bash, Ansible or Terraform
  • Bonus: NVIDIA certification, LLM training at scale, liquid-cooled GB300 NVL72 experience

If you'd rather be tuning NCCL collectives than sitting in another status meeting — this is your seat.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
GPU Cluster Architect
GPU Cluster Architect

Nebius B.V. • India

On-site
INR 4,000,000 - 6,500,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Corporation • India

On-site
INR 3,500,000 - 7,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Corporation • Mumbai

On-site
INR 3,500,000 - 5,000,000
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Bengaluru

On-site
INR 1,500,000 - 2,500,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000