Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
DAMAC Digital in Bengaluru is seeking an experienced HPC engineer to design and optimize GPU compute clusters for training and inference workloads. You will own architecture, deployment, and performance across bare metal, firmware, CUDA, NCCL, cuDNN, and the orchestration layer.
You'll integrate InfiniBand/RoCE fabrics, Slurm, Kubernetes and Run:AI, benchmark with NCCL tests and MLPerf, and lead vendor engagements with NVIDIA and partners.
Someone has to design and tune the GPU clusters that actually run DAMAC AI's training and inference workloads. That's this role.
We're building one of the region's most ambitious AI compute footprints — NVIDIA B200, B300 and GB300 NVL72 clusters powering sovereign cloud and hyperscale AI services across the Middle East and Asia. You'll own architecture, deployment and performance across the full stack: bare metal, firmware, CUDA, NCCL, orchestration and multi-tenant scheduling.
What you'll do
What you bring
If you'd rather be tuning NCCL collectives than sitting in another status meeting — this is your seat.