Head of Cluster Engineering

Nava

Singapore

On-site

SGD 250,000 - 420,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nava is building Asia’s next-generation AI-native cloud platform, seeking a Head of Cluster Engineering to lead design, orchestration, and scaling of multi-tenant accelerator fabrics for massive AI workloads. You’ll bridge deep technical expertise with strategic leadership to deliver low-latency inference and high-performance training across APAC hubs.

You will own the core systems powering our GPU clusters, drive Kubernetes/Slurm/Volcano orchestration, and partner with NVIDIA ecosystem teams to

Qualifications

  • 10+ years of experience in systems, infrastructure, or cluster engineering, with at least 5 years in a senior leadership role.
  • Deep expertise in Linux kernel internals, GPU driver stack (CUDA, UVM, persistence mode), and low-level firmware management (UEFI, BMC, IPMI).
  • Extensive experience with NVIDIA ecosystem tools and frameworks: CUDA, NCCL, NVLink/NVSwitch, DOCA, GPUDirect, TensorRT, and BlueField DPUs.
  • Proven success designing and scaling high-performance, multi-tenant GPU clusters for AI training/inference workloads at scale.

Responsibilities

  • Lead design and scaling of Mesa-scale AI cloud infrastructure and multi-tenant accelerator fabrics.
  • Oversee bare-metal provisioning, BIOS/UEFI configs, BMC/IPMI integrations, firmware pipelines, and host GPU driver stacks.

Skills

Linux kernel
GPU drivers
NVIDIA CUDA
NCCL
Kubernetes
Slurm
Volcano
InfiniBand
eBPF
Cilium
GitOps
Terraform
Telemetry

Tools

BIOS/UEFI
IPMI
BMC
GPU driver stack tooling
Terraform/Ansible

Job description

About Nava

Nava is building Asia’s next-generation AI-native cloud platform—designed to empower the world’s most ambitious AI developers with scalable, high-performance infrastructure. Our mission is to democratize access to enterprise-grade AI compute by delivering a secure, efficient, and developer-friendly cloud experience. While our technology is deeply rooted in systems innovation, we believe that great infrastructure should be invisible: enabling builders to focus on what matters most—their models, not the hardware.

About Nava

Nava is building Asia’s next-generation AI-native cloud platform—designed to empower the world’s most ambitious AI developers with scalable, high-performance infrastructure. Our mission is to democratize access to enterprise-grade AI compute by delivering a secure, efficient, and developer-friendly cloud experience. While our technology is deeply rooted in systems innovation, we believe that great infrastructure should be invisible: enabling builders to focus on what matters most—their models, not the hardware.

Role Responsibilities

We are rapidly scaling our AI-native GPU cloud to support massive-scale inference and training workloads across strategic Tier-1 APAC hubs. As we deploy multi-megawatt, high-density AI infrastructure, we’re looking for a Head of Cluster Engineering to lead the design, orchestration, driver/firmware integration, and scaling of our multi-tenant accelerator fabrics.

This is a critical leadership role where you’ll bridge deep technical expertise with strategic vision—ensuring our infrastructure delivers the performance, reliability, and flexibility needed to power next-generation AI applications. You’ll own the core systems that power our low-latency inference engine, high-performance training platforms, and the Nava Data Platform.

What You Will Do
  • Engineering Leadership: Build, scale, and mentor an elite engineering team dedicated to control plane architecture, deep systems tuning, driver/firmware automation, and cluster-scale orchestration.
  • Firmware & Driver Lifecycle: Direct automated bare-metal provisioning, BIOS/UEFI configurations, BMC/IPMI integrations, firmware flashing pipelines, and host GPU driver stack deployment across heterogeneous accelerator hardware.
  • NVIDIA Stack & Accelerator Integration: Drive deep hardware-software co-design across the NVIDIA ecosystem (CUDA, NVLink/NVSwitch, NCCL, GPUDirect Storage/RDMA, BlueField DPUs, DOCA, TensorRT) to maximize collective communication bandwidth and GPU utilization.
  • Advanced Orchestration: Drive the development and tuning of multi-tenant workload scheduling using Kubernetes, Slurm, and Volcano schedulers for topology-aware, fabric-conscious resource allocation.
  • Network & Interconnect Optimization: Optimize high-performance fabrics, leveraging InfiniBand, RoCE v2, RDMA, eBPF, and Cilium to eradicate network latency and tail congestion in large-scale distributed AI workloads.
  • Data Platform Interoperability: Seamlessly integrate infrastructure with the proprietary Nava Data Platform, tuning parallel file systems and distributed data pipelines for massive-scale training and low-latency inference.
  • Operational Excellence & Observability: Establish declarative GitOps infrastructure pipelines, robust SRE practices, and high-cardinality telemetry spanning hardware health, firmware status, thermals, interconnect error counters, and cluster performance metrics.
What We Are Looking For
  • Technical & Systems Mastery: Deep, hands-on architectural experience with large-scale bare-metal provisioning, Linux kernel internals, GPU host drivers, and low-level firmware management.
  • Ecosystem Expertise: Expert-level command of NVIDIA system architectures (NVLink topologies, NVSwitch fabrics, CUDA driver/runtime, NCCL tuning, DOCA/DPU offloading) alongside eBPF/Cilium networking and parallel storage systems.
  • Leadership Experience: Proven track record of building, managing, and scaling high-performing systems, kernel, network, or cluster engineering teams in production-critical environments.
  • Complex Problem Solving: Demonstrated ability to diagnose and solve complex cross-stack failures across driver/firmware boundary conditions, PCIe topologies, network transport layers, and distributed scheduler queues.
  • Execution Focus: A relentless drive for delivering highly available, deterministic, and scalable infrastructure capable of reliably backing bleeding-edge AI workloads.
Ideal Background

While experience at hyperscalers, GPU cloud providers, HPC organizations, or AI infrastructure startups is valuable, we welcome strong candidates from diverse technical backgrounds—including enterprise infrastructure, systems software, or even adjacent domains like fintech or robotics—if you can demonstrate deep systems thinking and a passion for AI infrastructure.

Qualifications
Required
  • 10+ years of experience in systems, infrastructure, or cluster engineering, with at least 5 years in a senior leadership role.
  • Deep expertise in Linux kernel internals, GPU driver stack (NVIDIA CUDA, UVM, persistence mode), and low-level firmware management (UEFI, BMC, IPMI).
  • Extensive experience with NVIDIA ecosystem tools and frameworks: CUDA, NCCL, NVLink/NVSwitch, DOCA, GPUDirect, TensorRT, and BlueField DPUs.
  • Proven success designing and scaling high-performance, multi-tenant GPU clusters for AI training/inference workloads at scale.
  • Strong background in orchestration systems: Kubernetes (K8s), Slurm, or Volcano schedulers—especially with topology-aware scheduling and fabric-aware resource allocation.
  • Hands-on experience optimizing high-speed interconnects: InfiniBand, RoCE v2, RDMA, and modern networking stacks (eBPF, Cilium).
  • Experience with GitOps, infrastructure-as-code (Terraform, Ansible), and observability tooling for hardware-level telemetry (e.g., health monitoring, error counters, thermals).
Preferred (Nice-to-Have)
  • Experience building or operating large-scale AI cloud platforms or GPU infrastructure at hyperscale providers (e.g., AWS, GCP, Azure, CoreWeave, Lambda Labs).
  • Familiarity with parallel file systems (e.g., Lustre, BeeGFS, WekaIO) and distributed data pipeline optimization for AI training.
  • Contributions to open-source projects related to kernel networking, GPU drivers, or cluster orchestration.
  • Background in HPC, high-frequency trading infrastructure, or cutting-edge AI research engineering.
The Opportunity:

This is a rare opportunity to lead the foundational cluster engineering strategy for a rapidly expanding AI cloud platform. You will have immense autonomy to shape our technological trajectory, build out an elite engineering team, and define next-generation GPU cloud infrastructure.

Why Join Us?

At Nava, you’ll be part of a mission-driven team building infrastructure for the next wave of AI innovation backed by top-tier investors and with deep roots in both systems engineering and product thinking. We value curiosity, ownership, and collaboration and we’re committed to creating an environment where engineers thrive, grow, and ship impactful technology.

Skills: training,kernel,nvidia,infrastructure,data,cloud,orchestration,cluster,cuda,firmware

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal / Distinguished Solution Architect
Principal / Distinguished Solution Architect

Nava • Singapore

On-site
SGD 180,000 - 300,000
AI-Cloud Cluster Engineering Lead
AI-Cloud Cluster Engineering Lead

Nava • Singapore

On-site
SGD 250,000 - 420,000
Solutions Architect (AI GPU Infrastructure / Data Centre Architecture)
Solutions Architect (AI GPU Infrastructure / Data Centre Architecture)

Nscale • Singapore

On-site
SGD 120,000 - 180,000
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 190,000
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SWAPETECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 170,000
Solutions Architect, Data Center MEP
Solutions Architect, Data Center MEP

NVIDIA Corporation • Singapore

On-site
SGD 150,000 - 190,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000