Principal Software Engineer, DGX Cloud Production Engineering

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe is seeking a Principal Software Engineer to shape the technical direction of our GPU infrastructure in Santa Clara, California. You will define the technical strategy for DGX Cloud cluster operations and lead the design and implementation of critical systems.

The ideal candidate has over 15 years of experience with distributed systems, deep knowledge of Kubernetes, and programming skills in Go or Python. A competitive salary ranging from $272,000 to $431,250 is offered, along with equity and benefits.

Qualifications

  • 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.
  • Deep experience with Kubernetes, Linux, infrastructure automation, and production operations.
  • Strong programming experience in Go, Python, or similar.

Responsibilities

  • Define and execute the technical strategy for DGX Cloud cluster operations.
  • Lead design and implementation of systems for cluster lifecycle, validation, and observability.
  • Establish patterns for Kubernetes-based GPU cluster operations.

Skills

Large-scale distributed systems
Kubernetes
Infrastructure automation
Go
Python

Education

BS/MS in Computer Science or equivalent experience

Job description

NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner, and cloud environments. We are looking for Principal Software Engineers to help shape the technical direction for production engineering, Kubernetes-based operations, automation, and reliability across large‑scale GPU clusters.

This role is for senior technical leaders who can define architecture, lead through influence, build critical systems, and turn ambiguous infrastructure problems into durable software and operating models.

What you’ll be doing:
  • Define and execute the technical strategy for DGX Cloud cluster operations, building the automation, GitOps, and Day 2 reliability needed to operate large‑scale GPU clusters across NVIDIA Cloud Partners (NCPs) and on‑prem environments.
  • Lead design and implementation of systems for cluster lifecycle, validation, repair, upgrades, observability, and readiness.
  • Establish patterns for Kubernetes‑based GPU cluster operations across partner and on‑prem environments.
  • Identify and eliminate operational toil through software, APIs, automation, and agent‑assisted workflows.
  • Set technical standards for production readiness, SLOs, incident response, handoff gates, and operational acceptance.
  • Mentor engineers and influence platform, infrastructure, storage, networking, security, and workload teams.
What we need to see:
  • 15+ years of experience building and operating large‑scale distributed systems or cloud infrastructure.
  • Deep experience with Kubernetes, Linux, infrastructure automation, and production operations.
  • Strong programming experience in Go, Python, or similar.
  • Proven ability to lead complex cross‑org technical initiatives.
  • Experience designing reliable systems with clear SLOs, observability, incident response, and automation.
  • BS/MS in Computer Science or equivalent experience.
Ways to stand out from the crowd:
  • Experience with GPU clusters, AI/ML infrastructure, Kubernetes operators, GitOps, BMaaS/VMaaS, managed Kubernetes, or multi‑cloud fleet operations.
  • Experience building internal platforms, control planes, lifecycle automation, or production readiness frameworks.
  • Track record of turning operational pain into reusable software, APIs, and engineering standards.

Base salary range: 272,000 USD – 431,250 USD. You will also be eligible for equity and benefits.

Applications will be accepted until May22,2026.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal‑opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, DGX Cloud Production Engineering
Principal Software Engineer, DGX Cloud Production Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

NVIDIA Corporation • Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

Socket.dev • Washington

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA • Washington

On-site
USD 272,000 - 431,000
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

Nvidia Corporation • Durham (NC)

On-site
USD 248,000 - 397,000
Equity
Benefits
Distinguished Engineer, Production Engineering, Data Center Automation
Distinguished Engineer, Production Engineering, Data Center Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 320,000 - 489,000
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA • Seattle (WA)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits package
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity