Head of GPU Cluster Engineering

Nava

Bengaluru

On-site

INR 6,000,000 - 11,000,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nava is building next-generation AI infrastructure and inference platforms. We are looking for a Head of GPU Cluster Engineering to lead architecture, deployment, and operations of large-scale GPU clusters.

This leadership role defines the engineering vision, drives execution across infrastructure domains, and ensures GPU platforms deliver world-class performance and reliability.

Qualifications

  • 12+ years in infrastructure engineering, distributed systems, cloud platforms, or AI infra.
  • Proven experience leading large-scale infrastructure or platform engineering teams.
  • Strong architectural thinking balancing technical excellence and business priorities.
  • Excellent stakeholder management and cross-functional leadership.

Responsibilities

  • Own end-to-end engineering outcomes from reference architecture to production operations.
  • Define vision, architecture, and operating model for large-scale GPU infrastructure.
  • Establish standards, design principles, and operational excellence across the platform.
  • Drive scalability, resiliency, and performance for growing AI workloads.
  • Lead cross-functional teams across GPU Compute, Networking, Storage, and SRE.

Skills

Leadership
Infrastructure engineering
GPU Compute
High-Speed Networking
Kubernetes
Distributed storage
Automation
Observability
Linux systems
Architectural thinking
Stakeholder management

Tools

CUDA
NCCL
Slurm
Ray
Kubeflow

Job description

About Nava

Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a


About Nava

Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a Head of GPU Cluster Engineering to lead the architecture, deployment, and operations of large-scale GPU clusters.


This is a highly strategic leadership role responsible for defining the engineering vision, driving execution across multiple infrastructure domains, and ensuring our GPU platforms deliver world-class performance, scalability, and reliability.


What You'll Do

Platform Leadership


  • Own end-to-end engineering outcomes, from reference architecture through production operations.

  • Define the technical vision, architecture, and operating model for large-scale GPU infrastructure.

  • Establish engineering standards, design principles, and operational excellence across the platform.

  • Drive platform scalability, resiliency, and performance to support rapidly growing AI workloads.


Architecture & Technical Strategy


  • Lead the design and evolution of GPU cluster architectures across compute, networking, storage, orchestration, and observability.

  • Evaluate and drive technology decisions around GPUs, interconnects, storage, Kubernetes, scheduling, and AI infrastructure software.

  • Review and approve architecture decisions while balancing performance, reliability, cost, and operational complexity.

  • Drive infrastructure standardization and automation across deployments.


Cross-Functional Leadership


  • Lead execution across multiple engineering pillars:


    • GPU Compute

    • High-Speed Networking

    • Storage

    • Platform Engineering

    • Site Reliability Engineering (SRE)

    • Infrastructure Automation


  • Partner closely with Product, Supply Chain, Data Centre Operations, and Customer Success teams to ensure successful platform delivery.

  • Act as the technical escalation point for critical engineering decisions.


Delivery & Operational Excellence


  • Own engineering readiness gates from design through production deployment.

  • Drive release planning, operational reviews, risk assessments, and post-incident analysis.

  • Establish SLAs, SLOs, and operational metrics for platform health and reliability.

  • Champion automation, observability, incident management, and continuous improvement across the engineering organization.


Team Leadership


  • Build, mentor, and scale a high-performing engineering organization.

  • Develop technical leaders across infrastructure disciplines.

  • Foster a culture of engineering excellence, ownership, collaboration, and innovation.

  • Support hiring and talent development for critical infrastructure roles.


What We're Looking For


  • 12+ years of experience in infrastructure engineering, distributed systems, cloud platforms, or AI infrastructure.

  • Proven experience leading large-scale infrastructure or platform engineering teams.

  • Deep understanding of GPU clusters, AI infrastructure, HPC, or large-scale distributed systems.

  • Strong expertise across:


    • GPU Compute (NVIDIA ecosystem preferred)

    • High-performance networking (InfiniBand, RoCE, RDMA)

    • Kubernetes and container orchestration

    • Distributed storage systems

    • Infrastructure automation and observability

    • Linux systems and platform engineering


  • Experience designing highly available, scalable production infrastructure.

  • Strong architectural thinking with the ability to balance technical excellence and business priorities.

  • Excellent stakeholder management and cross-functional leadership skills.


Nice to Have


  • Experience building AI factories, GPU cloud platforms, or large-scale inference infrastructure.

  • Familiarity with CUDA, NCCL, Slurm, Ray, Kubeflow, or similar AI infrastructure technologies.

  • Experience working with hyperscalers, cloud providers, or AI-first technology companies.

  • Exposure to multi-region or global infrastructure deployments.


Why Join Nava?


  • Build one of the world's leading AI infrastructure platforms.

  • Lead the engineering strategy behind next-generation GPU clusters and AI factories.

  • Work alongside world-class engineers solving some of the most challenging infrastructure problems in AI.

  • Shape the future of AI infrastructure from architecture through production at global scale.


Skills: leadership,automation,platforms,gpu,reliability,storage,architecture,cloud,infrastructure,cluster,drive

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Recruiter – Talent Acquisition & Employer Branding
Recruiter – Talent Acquisition & Employer Branding

Nava • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure

NVIDIA Gruppe • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Generous benefits package
Director, Site Reliability and Software Engineering DGX Cloud (Mumbai)
Director, Site Reliability and Software Engineering DGX Cloud (Mumbai)

NVIDIA • Mumbai

On-site
INR 17,143,000 - 26,667,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA Gruppe • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000