Principal Engineer – Cluster Deployment

Nava

Bengaluru

On-site

INR 4,200,000 - 6,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nava in Bengaluru seeks a Principal Engineer - Cluster Deployment to lead the end-to-end deployment of large-scale GPU clusters across global data centres. You will own deployments from hardware delivery to production-ready infrastructure, coordinating sites, vendors, and cross-functional teams to meet quality standards and customer readiness.

The role emphasizes on-site leadership, deep hardware lifecycle management, and driving playbooks, automation, and continuous improvement across

Qualifications

  • 10+ years in infrastructure deployment or HPC.
  • Experience deploying large-scale GPU clusters.
  • Server, rack deployment, high-speed networking, and power systems.
  • Experience coordinating across sites and vendors.
  • Willingness to travel domestically and internationally.
  • Familiarity with NVIDIA DGX/HGX and AI infrastructure.

Responsibilities

  • End-to-End Cluster Deployment: from rack arrival to tenant-ready production capacity.
  • Plan, coordinate, and execute deployment across multiple global data centres.
  • Drive installation, rack integration, power-up, firmware validation, and production readiness.
  • Ensure deployments meet timelines, quality standards, and safety requirements.
  • Lead cluster bring-up: hardware validation, network config, storage integration, platform initialization.
  • Oversee cabling, power validation, BIOS configuration, and infrastructure readiness.
  • Coordinate deployment activities across internal teams, hardware vendors and data centre partners.
  • Resolve deployment blockers and drive rapid issue resolution during implementation.
  • Develop and execute deployment validation procedures and acceptance criteria.
  • Lead burn-in testing, stress testing, hardware diagnostics, and production readiness validation.
  • Ensure clusters meet performance, stability, and reliability benchmarks before handover.
  • Own first-pass acceptance and minimize rework through robust quality processes.
  • Act as primary technical lead during on-site deployments.
  • Manage relationships with OEMs, contract manufacturers, data centre operators, and installation partners.
  • Ensure deployment standards are consistently followed across locations.
  • Drive continuous improvement across deployment processes, documentation, and execution methodologies.
  • Partner with GPU Cluster Engineering, Platform Engineering, Networking, SRE, Supply Chain, and Data Centre Operations teams.
  • Ensure seamless transition from deployment to production operations.
  • Support troubleshooting of deployment issues and coordinate engineering fixes where required.

Skills

Infrastructure deployment
GPU clusters
High-speed networking
Linux administration
Cross-functional leadership

Education

Bachelor's degree in CS/EE

Tools

Kubernetes

Job description

About Nava

Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a

About Nava

Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Principal Engineer - Cluster Deployment to lead the successful deployment of large-scale GPU clusters across global data centres.

This is a high-impact execution role responsible for taking GPU clusters from hardware delivery to fully validated, production-ready infrastructure. You'll own deployment execution across sites, vendors, and cross-functional teams, ensuring every cluster is delivered on time, meets quality standards, and is ready for customer workloads.

What You'll Do
  • End-to-End Cluster Deployment
    • Own the deployment of GPU clusters from rack arrival through to tenant-ready production capacity.
    • Plan, coordinate, and execute deployment activities across multiple global data centre locations.
    • Drive installation, rack integration, power-up, hardware bring-up, firmware validation, and production readiness.
    • Ensure deployments are completed safely, efficiently, and within agreed timelines.
  • Deployment Execution
    • Lead cluster bring-up activities, including hardware validation, network configuration, storage integration, and platform initialization.
    • Oversee structured cabling, power validation, firmware upgrades, BIOS configuration, and infrastructure readiness.
    • Coordinate deployment activities across internal engineering teams, hardware vendors, systems integrators, and data centre partners.
    • Resolve deployment blockers and drive rapid issue resolution during implementation.
  • Validation & Quality Assurance
    • Develop and execute deployment validation procedures and acceptance criteria.
    • Lead burn-in testing, stress testing, hardware diagnostics, and production readiness validation.
    • Ensure clusters meet defined performance, stability, and reliability benchmarks before customer handover.
    • Own first-pass acceptance and minimize deployment rework through robust quality processes.
  • Vendor & Site Management
    • Act as the primary technical lead during on-site deployments.
    • Manage relationships with OEMs, contract manufacturers, data centre operators, and installation partners.
    • Ensure deployment standards are consistently followed across all locations.
    • Drive continuous improvement across deployment processes, documentation, and execution methodologies.
  • Cross-Functional Collaboration
    • Partner with GPU Cluster Engineering, Platform Engineering, Networking, SRE, Supply Chain, and Data Centre Operations teams.
    • Ensure seamless transition from deployment to production operations.
    • Support troubleshooting of deployment issues and coordinate engineering fixes where required.
  • Operational Excellence
    • Develop deployment playbooks, SOPs, checklists, and automation to improve deployment consistency.
    • Track deployment metrics and identify opportunities to reduce deployment timelines and improve quality.
    • Drive lessons learned reviews following every deployment.
Success Metrics
You Will Be Measured On
  • Time-to-live for new GPU clusters
  • First-pass deployment acceptance rate
  • Deployment quality and reliability
  • Reduction in deployment rework
  • Deployment schedule adherence
  • Production readiness at handover
  • Deployment process standardization and automation
Qualifications
Required Qualifications
  • 10+ years of experience in infrastructure deployment, data centre engineering, HPC, cloud infrastructure, or systems engineering.
  • Proven experience deploying large-scale compute infrastructure, GPU clusters, HPC systems, or cloud platforms.
  • Deep technical expertise in:
    • Server and rack deployment
    • GPU hardware platforms
    • High-speed networking (InfiniBand, Ethernet, RDMA)
    • Structured cabling and power systems
    • Linux systems administration
    • Firmware, BIOS, and hardware lifecycle management
    • Infrastructure validation and production readiness
  • Experience coordinating cross-functional deployment projects across multiple sites.
  • Strong troubleshooting and problem-solving skills in complex infrastructure environments.
  • Willingness to travel to domestic and international deployment locations.
Preferred Qualifications
  • Experience with NVIDIA DGX, HGX, or similar GPU platforms.
  • Familiarity with Kubernetes, cluster provisioning, automation, and Infrastructure-as-Code.
  • Experience working with hyperscalers, AI infrastructure providers, or data centre operators.
  • Exposure to HPC, AI factories, or large-scale inference platforms.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of GPU Cluster Engineering
Head of GPU Cluster Engineering

Nava • Bengaluru

On-site
INR 6,000,000 - 11,000,000
Principal Engineer – GPU Orchestration
Principal Engineer – GPU Orchestration

Nava • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Field Technical Lead - Data Center Deployments
Field Technical Lead - Data Center Deployments

Nebius B.V. • Hyderabad

On-site
INR 900,000 - 1,500,000
Senior Project Delivery Manager, NVIS
Senior Project Delivery Manager, NVIS

NVIDIA Gruppe • Mumbai

On-site
INR 3,500,000 - 6,500,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior Project Delivery Manager, NVIS
Senior Project Delivery Manager, NVIS

NVIDIA • Mumbai

On-site
INR 3,500,000 - 6,000,000