Distinguished Engineer, Production Engineering, Data Center Automation - NVIDIA

OpenTalent

Santa Clara (CA)

On-site

USD 210,000 - 260,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA is seeking a Distinguished Engineer to lead cluster operations for DGX Cloud capacity across on-prem and cloud environments.

The role blends software engineering, systems knowledge, and production discipline to build reliable, scalable DGX Cloud platforms and workflows that serve researchers and customers.

You will define technical strategy, set operating standards, and drive cross-team delivery to improve production readiness and overall platform availability at scale.

Responsibilities

  • Define the long-range technical strategy for operating DGX Cloud clusters across on-prem, hyperscalers, and NeoCloud environments.
  • Define the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability.
  • Guide the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination.
  • Make and guide high-impact technical decisions that resolve how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production.

Job description

NVIDIA is looking for a Distinguished Engineer to act as a senior technical leader in the Production Engineering group, enthusiastic about cluster operations involving DGX Cloud GPU capacity.

At NVIDIA, Production Engineering is responsible for ensuring large-scale production systems are reliable, straightforward to lead, and increasingly automated across NVIDIA's DGX Cloud resources. We combine software engineering, systems engineering, and extensive production knowledge to build platforms, workflows, and operational frameworks that sustain GPU infrastructure health, scalability, and availability for researchers and customers.

This role centers on the operational framework supporting DGX Cloud environments spanning on-premises, major cloud providers, and NVIDIA Cloud Partner locations. The responsibilities include engineering integrations to guarantee DGX Cloud capacity is fully operational in production. This involves Kubernetes service management, ensuring vendor and equipment availability, on-prem infrastructure operations, release and runtime preparation, and service reliability coordination. The workflows tie these elements into a cohesive production system.

This hands-on Distinguished Engineer role calls for a deeply technical leader to build the architectural direction for cluster operations in DGX Cloud. The ideal candidate will blend software engineering expertise, system knowledge, and production insight to define technical strategy, set operating standards, direct the evolution of the production model, and drive delivery of cross-organizational capabilities. These capabilities ensure that the DGX Cloud resources remain usable, maintainable, and improve continuously at scale. The position demands both deep invention and implementation skills and the ability to lead by influence across several teams and critical production results.

What you'll be doing:

Define the long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments

Define the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability throughout DGX Cloud resources

Guide the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination

Make and guide high-impact technical decisions that resolve how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production

Develop robust workflows, interfaces, and engineering collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability domains

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Engineer, Production Engineering, Cluster Management - nvidia
Distinguished Engineer, Production Engineering, Cluster Management - nvidia

OpenTalent • Santa Clara (CA)

On-site
USD 250,000 - 420,000
Distinguished Engineer, Production Engineering, Data Center Automation
Distinguished Engineer, Production Engineering, Data Center Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 320,000 - 489,000
Distinguished Engineer, DGX Cloud Production & Automation
Distinguished Engineer, DGX Cloud Production & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 320,000 - 489,000
Distinguished Engineer, DGX Cloud Cluster Operations
Distinguished Engineer, DGX Cloud Cluster Operations

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Senior Distinguished Engineer, DGX Cloud Cluster Ops
Senior Distinguished Engineer, DGX Cloud Cluster Ops

Socket.dev • Santa Clara (CA)

Hybrid
USD 320,000 - 489,000
Equity
Benefits
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Distinguished Engineer, DGX Cloud Production & Automation
Distinguished Engineer, DGX Cloud Production & Automation

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

Socket.dev • Santa Clara (CA)

Hybrid
USD 320,000 - 489,000
Equity
Benefits
DGX Cloud Production Lead & Clusters Architect
DGX Cloud Production Lead & Clusters Architect

OpenTalent • Santa Clara (CA)

On-site
USD 250,000 - 420,000
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity