Principal Software Engineer, DGX Cloud Production Engineering

NVIDIA

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Principal Software Engineer to define architecture and lead operations of large-scale GPU clusters. The role requires strong expertise in Kubernetes, Linux, and infrastructure automation.

Candidates should have over 15 years of experience with distributed systems, proven leadership in technical initiatives, and a degree in Computer Science. A competitive salary between $272,000 and $431,250 based on experience and location will be offered.

Qualifications

  • 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.
  • Deep experience with Kubernetes and production operations.
  • Strong programming experience in Go, Python, or similar.

Responsibilities

  • Define and execute the technical strategy for DGX Cloud cluster operations.
  • Lead design and implementation of systems for cluster lifecycle and readiness.
  • Identify and eliminate operational toil through software and automation.

Skills

Kubernetes
Linux
Infrastructure automation
Go
Python
Cloud infrastructure

Education

BS/MS in Computer Science or equivalent experience

Job description

NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner, and cloud environments. We are looking for Principal Software Engineers to help shape the technical direction for production engineering, Kubernetes-based operations, automation, and reliability across large-scale GPU clusters.

This role is for senior technical leaders who can define architecture, lead through influence, build critical systems, and turn ambiguous infrastructure problems into durable software and operating models.

What you’ll be doing:
  • Define and execute the technical strategy for DGX Cloud cluster operations, building the automation, GitOps, and Day 2 reliability needed to operate large-scale GPU clusters across NVIDIA Cloud Partners (NCPs) and on-prem environments.

  • Lead design and implementation of systems for cluster lifecycle, validation, repair, upgrades, observability, and readiness.

  • Establish patterns for Kubernetes-based GPU cluster operations across partner and on-prem environments.

  • Identify and eliminate operational toil through software, APIs, automation, and agent-assisted workflows.

  • Set technical standards for production readiness, SLOs, incident response, handoff gates, and operational acceptance.

  • Mentor engineers and influence platform, infrastructure, storage, networking, security, and workload teams.

What we need to see:
  • 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.

  • Deep experience with Kubernetes, Linux, infrastructure automation, and production operations.

  • Strong programming experience in Go, Python, or similar.

  • Proven ability to lead complex cross-org technical initiatives.

  • Experience designing reliable systems with clear SLOs, observability, incident response, and automation.

  • BS/MS in Computer Science or equivalent experience.

Ways to stand out from the crowd:
  • Experience with GPU clusters, AI/ML infrastructure, Kubernetes operators, GitOps, BMaaS/VMaaS, managed Kubernetes, or multi-cloud fleet operations.

  • Experience building internal platforms, control planes, lifecycle automation, or production readiness frameworks.

  • Track record of turning operational pain into reusable software, APIs, and engineering standards.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 22, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Principal Software Engineer, DGX Cloud Production Engineering
Principal Software Engineer, DGX Cloud Production Engineering

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA • Seattle (WA)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits package
Senior Software Engineer, DGX Cloud Production Engineering
Senior Software Engineer, DGX Cloud Production Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

2100 NVIDIA USA • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Comprehensive benefits package
Equity options
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior Software Engineer, DGX Cloud Production Engineering
Senior Software Engineer, DGX Cloud Production Engineering

Segment (Twilio) • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Equity
Benefits
Principal Software Engineer - DGX Cloud
Principal Software Engineer - DGX Cloud

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package