AI-Cloud Cluster Engineering Lead

Nava

Singapore

On-site

SGD 250,000 - 420,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nava is building Asia’s next-generation AI-native cloud platform, seeking a Head of Cluster Engineering to lead design, orchestration, and scaling of multi-tenant accelerator fabrics for massive AI workloads. You’ll bridge deep technical expertise with strategic leadership to deliver low-latency inference and high-performance training across APAC hubs.

You will own the core systems powering our GPU clusters, drive Kubernetes/Slurm/Volcano orchestration, and partner with NVIDIA ecosystem teams to

Qualifications

  • 10+ years of experience in systems, infrastructure, or cluster engineering, with at least 5 years in a senior leadership role.
  • Deep expertise in Linux kernel internals, GPU driver stack (CUDA, UVM, persistence mode), and low-level firmware management (UEFI, BMC, IPMI).
  • Extensive experience with NVIDIA ecosystem tools and frameworks: CUDA, NCCL, NVLink/NVSwitch, DOCA, GPUDirect, TensorRT, and BlueField DPUs.
  • Proven success designing and scaling high-performance, multi-tenant GPU clusters for AI training/inference workloads at scale.

Responsibilities

  • Lead design and scaling of Mesa-scale AI cloud infrastructure and multi-tenant accelerator fabrics.
  • Oversee bare-metal provisioning, BIOS/UEFI configs, BMC/IPMI integrations, firmware pipelines, and host GPU driver stacks.

Skills

Linux kernel
GPU drivers
NVIDIA CUDA
NCCL
Kubernetes
Slurm
Volcano
InfiniBand
eBPF
Cilium
GitOps
Terraform
Telemetry

Tools

BIOS/UEFI
IPMI
BMC
GPU driver stack tooling
Terraform/Ansible

Job description

Nava is building Asia’s next-generation AI-native cloud platform, seeking a Head of Cluster Engineering to lead design, orchestration, and scaling of multi-tenant accelerator fabrics for massive AI workloads. You’ll bridge deep technical expertise with strategic leadership to deliver low-latency inference and high-performance training across APAC hubs.

You will own the core systems powering our GPU clusters, drive Kubernetes/Slurm/Volcano orchestration, and partner with NVIDIA ecosystem teams to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Cluster Engineering
Head of Cluster Engineering

Nava • Singapore

On-site
SGD 250,000 - 420,000
APAC AI Infrastructure Architect & Strategy Lead
APAC AI Infrastructure Architect & Strategy Lead

Nava • Singapore

On-site
SGD 180,000 - 300,000
Principal / Distinguished Solution Architect
Principal / Distinguished Solution Architect

Nava • Singapore

On-site
SGD 180,000 - 300,000
Senior AI/HPC Compute Architect for GPU Clusters
Senior AI/HPC Compute Architect for GPU Clusters

NVIDIA • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Infra Architect: GPU Clusters & HPC Ops
Senior AI Infra Architect: GPU Clusters & HPC Ops

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: GPU HPC Clusters & Orchestration
AI Infra Engineer: GPU HPC Clusters & Orchestration

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Infra Engineer — GPU Cluster & ML Platform
Senior AI Infra Engineer — GPU Cluster & ML Platform

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
APAC AI/HPC Delivery Leader — InfiniBand & DC
APAC AI/HPC Delivery Leader — InfiniBand & DC

NVIDIA Gruppe • Singapore

On-site
SGD 150,000 - 200,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
Senior AI/HPC Compute Infrastructure Engineer
Senior AI/HPC Compute Infrastructure Engineer

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 190,000