Senior Software Engineer, Kubernetes Runtime and Release

NVIDIA Corporation

California, Northern (MO, KY)

Hybrid

USD 184,000 - 357,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation is hiring a Senior Software Engineer for the Kubernetes Runtime and Release team. You will work on GPU cluster software, building Go controllers and APIs to install, upgrade, and validate systems across clouds and hardware providers.

The role emphasizes Kubernetes expertise, controllers and release automation, and collaboration across teams. Opportunities include involvement in production-grade pipelines and early access to new silicon in a fast-growing cloud infra project.

Qualifications

  • 6+ years building production infrastructure software or distributed systems.
  • Strong Go programming and Kubernetes experience, with focus on controllers or release automation.

Responsibilities

  • Build Go controllers and APIs to install, upgrade, and validate GPU cluster software.
  • Develop release qualification pipelines and automation across providers.

Skills

Go programming
Kubernetes expertise
Distributed systems
Go controllers

Education

BS in Computer Science
MS in Computer Science

Tools

controller-runtime
CRDs
Reconcilers

Job description

## Senior Software Engineer, Kubernetes Runtime and ReleaseApply: US, CA, Remote: Full time: Posted Today: JR2026573NVIDIA researchers depend on GPU clusters for large-scale AI workloads. Our DGX Cloud Kubernetes Runtime & Release team brings those clusters to life across major public clouds and specialized GPU providers, often on hardware that is new to the world when we get it. We build and maintain the supported Kubernetes runtime, automate its delivery, and bring new providers and GPU platforms into production. We’re growing quickly and taking on broader ownership of NVIDIA’s cluster software delivery. We’re hiring across Runtime, Release Engineering, and Provider Integration, with each role focused on your strengths. You don’t need experience across every area below. **What you’ll be doing:**Your primary focus will be one of three areas, with collaboration across the team:* Runtime: Build Go controllers and APIs to install, upgrade, and validate GPU cluster software. Integrate components, define API contracts, and evolve Helm and Argo CD delivery toward controller-driven automation.* Release Engineering: Build validation pipelines that inform release decisions across providers and GPU platforms. Develop systems to allocate GPU capacity across validation runs and account for cloud reservations and quotas. Make qualification more efficient through reusable tests and clear failure reports.* Provider Integration: Bring new providers and GPU hardware into production, potentially among the first engineers working with new silicon. Resolve integration failures with partner teams and turn initial provisioning, upgrade, and operational checks into repeatable automation. ## **What we need to see:*** 6+ years building production infrastructure software or distributed systems.* Strong programming skills in Go or another language to build production systems, with willingness to work primarily in Go.* Kubernetes experience and depth in at least one area: controllers and operators, release automation, test and validation systems, or cloud integration.* Experience delivering engineering projects, diagnosing complex failures, and collaborating across teams.* BS or MS in Computer Science, Engineering, or equivalent experience. **Ways to stand out from the crowd:*** Experience in any of these areas is valuable, but not required:* Go development with controller-runtime, CRDs, and reconcilers.* Release qualification across multiple environments or platforms.* GPU infrastructure, accelerated networking, or GPU scheduling.* Bringing new hardware, regions, or cloud providers into production.* Resource allocation, leasing, or fair-share scheduling and upstream integration, compatibility, or software supply chain integrity. This role suits an engineer who wants direct influence over what reaches production, and who builds for the hundredth cluster while shipping the first. Join us and help build the next generation of NVIDIA’s GPU cloud infrastructure!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until October 3, 2026.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Kubernetes Runtime and Release
Senior Software Engineer, Kubernetes Runtime and Release

NVIDIA Corporation • Kansas

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer, Kubernetes Runtime and Release
Senior Software Engineer, Kubernetes Runtime and Release

NVIDIA • United States

Remote
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer, DGX Cloud Production Engineering
Senior Software Engineer, DGX Cloud Production Engineering

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 224,000 - 357,000
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 184,000 - 288,000
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior Manager, Kubernetes Runtime Engineering
Senior Manager, Kubernetes Runtime Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Comprehensive benefits
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Software Engineer
Senior Software Engineer

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 272,000 - 431,000
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits