Kubernetes Runtime Engineering Lead — Multi-Tenant GPU Platform

NVIDIA

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA's Kubernetes Engine (NKE) team is seeking a technical leader to guide the Runtime Engineering group responsible for the full configuration lifecycle of NKE tenant workload clusters.

You will oversee the build, deployment, and operational reliability of cluster configurations across topologies, manage a team of engineers across AICR, DCGM, and related components, and drive architecture decisions for networking, storage, and security to deliver a production-grade, multi-tenant Kubernetes

Qualifications

  • BS/MS in CS or related field (or equivalent experience).
  • 12+ years designing and delivering large-scale distributed software systems; 5+ years in people management.
  • Experience bridging runtime, networking, and security fields.
  • Kubernetes internals knowledge beyond usage (scheduler, kubelet, API server, admission).
  • Cluster lifecycle management experience (Cluster API, kubeadm, etc.).
  • Security and compliance posture (CIS Kubernetes Benchmark, image signing, supply chain).

Responsibilities

  • Oversee build, implementation, and operational reliability of cluster configurations for NKE tenant workloads across all topologies.
  • Lead engineers coordinating the container runtime stack (AICR, GPU management operator, DCGM, and related components).
  • Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster HA, and GPU resource partitioning (MIG, MPS, time-slicing).
  • Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries.
  • Collaborate with NKE platform, infrastructure, and cybersecurity teams to integrate capabilities and resolve runtime concerns.
  • Build tooling for AICR lifecycle management — provisioning, upgrades, drift detection, remediation.
  • Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership.
  • Contribute to open source communities wherever NKE has upstream dependencies or influence.

Skills

Team leadership
Distributed systems
Kubernetes internals
Cluster lifecycle management
Security & compliance
API design
IAM approaches
Cross-org collaboration

Education

BS/MS in CS or related field

Tools

Kubernetes
Cluster API/kubeadm
AICR
DCGM

Job description

NVIDIA's Kubernetes Engine (NKE) team is seeking a technical leader to guide the Runtime Engineering group responsible for the full configuration lifecycle of NKE tenant workload clusters.

You will oversee the build, deployment, and operational reliability of cluster configurations across topologies, manage a team of engineers across AICR, DCGM, and related components, and drive architecture decisions for networking, storage, and security to deliver a production-grade, multi-tenant Kubernetes

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Kubernetes Runtime Lead - GPU AI Platform
Senior Kubernetes Runtime Lead - GPU AI Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Comprehensive benefits
Senior Manager, Kubernetes Runtime Engineering
Senior Manager, Kubernetes Runtime Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Comprehensive benefits
Senior Manager, Kubernetes Runtime Engineering
Senior Manager, Kubernetes Runtime Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Senior Kubernetes Runtime & Release Engineer
Senior Kubernetes Runtime & Release Engineer

NVIDIA • United States

Remote
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Runtime & Release Engineer - GPU Cloud
Senior Kubernetes Runtime & Release Engineer - GPU Cloud

NVIDIA Corporation • California (MO), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Runtime & Release Engineer — GPU Cloud
Senior Kubernetes Runtime & Release Engineer — GPU Cloud

NVIDIA Corporation • Kansas

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Node Lifecycle Architect
Senior Kubernetes Node Lifecycle Architect

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Lead, GPU Kubernetes Runtime & Open-Source Strategy
Lead, GPU Kubernetes Runtime & Open-Source Strategy

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 208,000 - 380,000
Equity
Benefits package
Senior Cloud Platform Engineer - Kubernetes & GitOps
Senior Cloud Platform Engineer - Kubernetes & GitOps

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits