Senior Manager, Kubernetes Runtime Engineering

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Comprehensive benefits

Job summary

NVIDIA Corporation is seeking a technical leader to head the Runtime Engineering team for the NVIDIA Kubernetes Engine (NKE). You will drive the lifecycle of tenant workload clusters, ensure reliability, scalability, and security across the platform, and coordinate with cross-functional teams to deliver a production-grade Kubernetes runtime.

You will manage architecture decisions for networking, storage, and GPU resource partitioning, and contribute to open sources.

Qualifications

  • 12+ years in designing and delivering large-scale distributed software systems.
  • 5+ years of people-management leading software teams.
  • Experience bridging runtime, networking, and security across orgs.

Responsibilities

  • Oversee build, implementation and reliability of cluster configurations for NKE tenant workloads.
  • Lead a team coordinating the container runtime stack: AICR, GPU management operator, DCGM.
  • Drive architecture decisions for cluster networking, storage, and GPU resource partitioning.
  • Define cluster hardening standards, RBAC models, and multi-tenancy boundaries.
  • Collaborate with platform, infra, and cybersecurity teams to integrate capabilities.
  • Build tooling for AICR lifecycle management: provisioning, upgrades, drift detection.

Skills

Leadership
Kubernetes
Distributed systems
Security & compliance
RBAC & pod security
API design
Cross-org collaboration

Education

BS/MS in Computer Science or related field

Tools

Cluster API
kubeadm
NVIDIA DCGM/AI Container Runtime (AICR)

Job description

The NVIDIA Kubernetes Engine (NKE) team is looking for a technical leader to lead the Runtime Engineering team responsible for the full configuration lifecycle of NKE tenant workload clusters. This team is responsible for software components that keep GPU workloads reliable and secure at scale. Their scope includes cluster bootstrapping, node configuration, and the container execution environment, including NVIDIA's AI Container Runtime (AICR). You will work across networking, storage, GPU resource management, and cluster security to deliver a production-grade, multi-tenant Kubernetes platform. Your team's decisions directly shape the runtime foundation that internal and external customers depend on.

What You'll Be Doing:
  • Be responsible for the build, implementation, and operational reliability of cluster configurations for NKE tenant workloads across all supported topologies
  • Manage a team of engineers coordinating the entire container runtime stack: AICR, GPU management operator, DCGM, and related node-level components
  • Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster HA , and GPU resource partitioning (MIG, MPS, time-slicing)
  • Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries
  • Partner with NKE platform, infrastructure, and cybersecurity teams to integrate new capabilities and resolve cross-cutting runtime concerns
  • Build and maintain tooling for AICR lifecycle management — provisioning, upgrades, configuration drift detection, and remediation
  • Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership
  • Contribute to open source communities anywhere NKE has upstream dependencies or influence
What We Need to See:
  • BS/MS degree in Computer Science or related field (or equivalent experience) 12+ overall years of relevant experience designing and delivering large-scale distributed software systems, including 5+ years of people-management experience leading, developing, and scaling high-performing software engineering teams responsible for complex, production-critical software.
  • Experience leading a group of engineers with varying specializations and seniority levels — bridging runtime, networking, and security fields is a core part of this role
  • Kubernetes internals knowledge — not just usage; you understand how the scheduler, kubelet, API server, and admission controllers interact
  • Cluster lifecycle management experience — Cluster API, kubeadm, or equivalent; experience leading fleet-scale cluster provisioning and upgrades
  • Security and compliance posture — CIS Kubernetes Benchmark, pod security admission, image signing, supply chain integrity
  • Proven ability to design and implement maintainable APIs for consumers
  • Familiarity with Identity and Access Management approaches
  • Excel in managing up, down, and across organizations
  • Demonstrated ability to reach cross-organization consensus without all the details
Ways to Stand Out from the crowd:
  • Prior experience with NVIDIA GPU Operator, DCGM Exporter, or NVLink-aware scheduling
  • Experience running Kubernetes at hyperscale with GPU node pools
  • Track record of upstream open source contributions in the Kubernetes or any open source runtime ecosystem
  • Experienced, persuasive, and effective interpersonal skills — written, verbal, and in front of engineering leadership
  • Demonstrated skills in coaching, analysis, problem solving, and short/long-term technical planning

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction — from artificial intelligence to autonomous vehicles.

NVIDIA is widely considered one of the technology world's most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. If you're passionate about building the infrastructure that runs AI at scale, we want to hear from you.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/ Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Kubernetes Runtime and Release
Senior Software Engineer, Kubernetes Runtime and Release

NVIDIA Corporation • Kansas

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Product Manager, Kubernetes Accelerated Runtime
Senior Product Manager, Kubernetes Accelerated Runtime

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 380,000
Equity
Benefits package
Senior Product Manager, Kubernetes Accelerated Runtime
Senior Product Manager, Kubernetes Accelerated Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 380,000
Senior Product Manager, Kubernetes Accelerated Runtime
Senior Product Manager, Kubernetes Accelerated Runtime

NVIDIA • Seattle (WA)

On-site
USD 208,000 - 380,000
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 184,000 - 288,000
Senior Software Engineer
Senior Software Engineer

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale – DGX Cloud
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale – DGX Cloud

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1