Principal Engineer – GPU Orchestration

Nava

Bengaluru

On-site

INR 4,000,000 - 8,000,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nava is building next-generation AI infrastructure and inference platforms powering enterprise AI at scale. We seek a Principal Engineer – GPU Orchestration to lead the design and evolution of our GPU orchestration platform for inference workloads on small GPU clusters.

You will own the orchestration layer, optimize utilization, enable multi-tenant isolation, and drive automation across Kubernetes and platform engineering teams, shaping reliable, scalable AI infrastructure.

Qualifications

  • 10+ years of experience in distributed systems, cloud infrastructure, Kubernetes, or platform engineering.
  • Deep expertise in Kubernetes internals, scheduling, and container orchestration.
  • Strong understanding of GPU resource management and AI infrastructure.
  • Experience building large-scale scheduling and orchestration systems.
  • Proficiency in Go, Python, or C++ for systems programming.
  • Excellent problem-solving, architecture, and technical leadership skills.

Responsibilities

  • Own architecture and implementation of GPU scheduling for inference workloads and small GPU clusters.
  • Design intelligent workload scheduling algorithms maximizing GPU utilization with fairness.
  • Manage GPU allocation, quotas, and resource sharing across customers.
  • Optimize scheduling policies to improve efficiency and reduce costs.
  • Own Kubernetes orchestration layer for AI workloads and cluster lifecycle.

Skills

GPU orchestration
Kubernetes
Scheduling
Distributed systems
Go
Python
C++

Tools

CUDA
MIG
GPU Operator
Kubernetes device plugins

Job description

About Nava

Nava is building next-generation AI infrastructure and inference platforms that power enterprise AI at scale. We are looking for a Principal Engineer – GPU Orchestration to lead the design and evolution of our GPU orchestration platform, enabling efficient scheduling, resource sharing, and multi-tenant inference workloads. This role will own the orchestration layer that manages GPU resources for inference and small GPU clusters (2–4 node deployments), ensuring optimal utilization, rapid scaling, workload isolation, and an exceptional customer experience.

What You'll Do
GPU Orchestration & Scheduling
  • Own the architecture and implementation of GPU scheduling for inference workloads and small GPU clusters (2–4 node deployments).
  • Design intelligent workload scheduling algorithms that maximize GPU utilization while ensuring fairness and predictable performance.
  • Manage GPU allocation, quotas, and resource sharing across multiple customers and workloads.
  • Continuously optimize scheduling policies to improve efficiency and reduce infrastructure costs.
Kubernetes & Platform Engineering
  • Own the Kubernetes orchestration layer for AI workloads.
  • Design and maintain GPU-aware scheduling capabilities within Kubernetes.
  • Improve cluster lifecycle management, workload placement, and infrastructure automation.
  • Partner with Platform Engineering teams to continuously enhance cluster reliability and scalability.
Multi-Tenancy & Resource Isolation
  • Design secure multi-tenant GPU environments with strong workload isolation.
  • Implement resource quotas, admission controls, namespace isolation, and scheduling policies.
  • Ensure consistent customer experience while maintaining high infrastructure utilization.
  • Collaborate closely with the Security team on tenancy and access control mechanisms.
Autoscaling & Workload Optimization
  • Build fast autoscaling capabilities for inference workloads.
  • Reduce cold-start times through intelligent provisioning and pre-warming strategies.
  • Optimize model placement based on workload characteristics, GPU availability, and latency requirements.
  • Improve workload elasticity while balancing cost and performance.
Performance Engineering
  • Define and monitor platform KPIs including:
    • GPU utilization
    • Scheduling latency
    • Cluster efficiency
    • Autoscaling performance
    • Cold-start latency
    • Workload throughput
  • Drive continuous improvements through benchmarking, performance tuning, and automation.
Cross-Functional Collaboration
  • Partner with Compute & Inference Platform, GPU Cluster Engineering, Platform Reliability, AI Infrastructure Security, and Product teams.
  • Support onboarding of new AI models and customer workloads.
  • Contribute to platform architecture decisions across Nava's AI infrastructure.
Technical Leadership
  • Serve as the technical authority for GPU orchestration and workload scheduling.
  • Mentor senior engineers and contribute to engineering best practices.
  • Drive architectural reviews, technical design discussions, and long-term platform strategy.
Success Metrics
You Will Be Measured On
  • GPU utilization across the platform
  • Scheduling efficiency and fairness
  • Autoscaling responsiveness
  • Cold-start latency
  • Multi-tenant performance and isolation
  • Platform reliability and scalability
  • Customer workload performance
  • Infrastructure cost optimization
Qualifications
Required Qualifications
  • 10+ years of experience in distributed systems, cloud infrastructure, Kubernetes, or platform engineering.
  • Deep expertise in Kubernetes internals, scheduling, and container orchestration.
  • Strong understanding of GPU resource management and AI infrastructure.
  • Experience building large-scale scheduling and orchestration systems.
  • Expertise in:
    • Kubernetes
    • Container runtimes
    • Distributed systems
    • Infrastructure automation
    • Resource scheduling
    • Multi-tenant platform architecture
    • Performance optimization
  • Strong software engineering skills in Go, Python, C++, or similar systems programming languages.
  • Excellent problem-solving, architecture, and technical leadership skills.
Preferred Qualifications
  • Experience with NVIDIA GPUs, CUDA, MIG (Multi-Instance GPU), GPU Operator, or Kubernetes device plugins.
  • Familiarity with KServe, Ray Serve, Triton Inference Server, vLLM, Slurm, or similar AI infrastructure technologies.
  • Experience with large-scale inference platforms, GPU cloud providers, or HPC environments.
  • Knowledge of Kubernetes scheduler extensions, custom controllers, and operator development.
Why Join Nava?
  • Build the orchestration platform powering one of the world's leading AI inference platforms.
  • Solve challenging problems in GPU scheduling, resource optimization, and distributed systems.
  • Work with world-class engineers building cutting-edge AI infrastructure.
  • Shape the future of enterprise AI by enabling efficient, scalable, and secure GPU orchestration at global scale.

Skills: customer,architecture,infrastructure,gpu,design,utilization,kubernetes,cluster,orchestration,scheduling,isolation

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of GPU Cluster Engineering
Head of GPU Cluster Engineering

Nava • Bengaluru

On-site
INR 6,000,000 - 11,000,000
Head of Compute & Inference Platform
Head of Compute & Inference Platform

Nava • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure

NVIDIA Gruppe • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Generous benefits package
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA Gruppe • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Gruppe • Pune District

On-site
INR 400,000 - 660,000
Competitive salary
Generous benefits package
Senior CAD Engineer
Senior CAD Engineer

NVIDIA Gruppe • Bengaluru

Hybrid
INR 3,500,000 - 5,200,000
Senior CAD Engineer
Senior CAD Engineer

NVIDIA • Hyderabad

Hybrid
INR 1,800,000 - 3,200,000