Distinguished Engineer, Production Engineering, Cluster Management

Jobtailor

California (MO)

On-site

USD 210,000 - 320,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks a senior technical leader to define and execute the strategy for operating DGX Cloud clusters across multiple environments. You will direct architectural direction, reliability, and operational standards to enable scalable, secure production deployments.

The role requires deep expertise in Kubernetes-based production systems, distributed software, and automation, with a track record of cross-organizational leadership and measurable outcomes.

Qualifications

  • 18+ years of experience building and operating large-scale distributed systems or production environments.
  • Proven track record leading cross-organizational technical efforts from concept to production.
  • Experience defining operating models, architectural direction, and engineering standards.

Responsibilities

  • Define long-range technical strategy for operating DGX Cloud clusters across data centers and hyperscalers.
  • Set architectural direction and standards for cluster lifecycle, runtime delivery, and release readiness.
  • Guide roadmaps and cross-functional investments to improve production readiness and performance.
  • Lead high-impact technical decisions with platform, hardware, provider, and service teams.
  • Build automation, APIs, and workflows to move capacity into production and balance existing capacity.

Skills

Kubernetes
InfrastructureAutomation
DistributedSystems
Python/Go
TechnicalLeadership
ProductionReliability

Education

BS/MS/PhD in CS or related field

Tools

Kubernetes
Cloud Platforms
Automation Tools

Job description

  • Define the long‑range technical strategy for operating DGX Cloud clusters consistently across local data centers, hyperscalers, and NeoCloud environments
  • Define architectural direction and operating standards for cluster lifecycle, runtime delivery, restoration, release readiness, and steady‑state operability
  • Guide roadmaps and execution of cross‑organizational investments improving production readiness, operational safety, performance, and coordination
  • Make and influence high‑impact technical decisions across platform, hardware, provider, and service teams
  • Build workflows, interfaces, and engineering handshakes across Kubernetes production service, provider and hardware preparation, on‑prem, and bare‑metal operations
  • Restructure operations and service‑layer reliability domains
  • Build and evolve automation, APIs, operating workflows, and readiness gates for moving new capacity into stable production and balancing existing capacity
  • Implement production operating approaches that reduce manual input and increase ownership clarity, consistency, traceability, and release safety
  • Partner with platform, hardware, provider engineering, service owners, and Production Engineering leaders to convert recurring friction into durable improvements
  • Raise standards for operability, resilience, scalability, and performance through build leadership, architecture review, and technical standards
Requirements
  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience
  • 18+ years of experience building and operating large‑scale distributed systems, infrastructure platforms, or production environments
  • Company‑level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms
  • Track record of defining operating models, architectural direction, and engineering standards across multiple technical domains and organizations
  • Record leading large, cross‑team technical efforts from concept through production, including aligning collaborators, navigating for clarity, and delivering measurable outcomes
  • Deep experience with Kubernetes‑based production systems, infrastructure automation, or distributed systems operations
  • Strong software engineering skills in Python, Go, or similar low‑level programming languages
  • Deep understanding of distributed systems, Linux, networking, containers, and production reliability concerns
  • Experience crafting operational workflows, APIs, service interfaces, or automation frameworks that become the standard way teams run production systems
  • Strong architectural judgment and a validated history of simplifying complex operational problems through reusable software, clear technical strategy, and durable engineering direction
Core Competencies

Demonstrates expertise in defining technical strategies for large‑scale distributed systems and cloud platforms, with a focus on operational excellence, architectural direction, and automation. Proven ability to lead cross‑organizational technical efforts and improve production readiness through effective collaboration and engineering standards.

Highest-signal resume keywords
  • Kubernetes‑Based Production Systems
  • Infrastructure Automation
  • Distributed Systems Operations
  • Software Engineering in Python or Go
  • Technical Leadership in Production Engineering
Hard Skills
  • Distributed Systems
  • Infrastructure Platforms
  • Production Environments
  • Operational Workflows
  • APIs
  • Service Interfaces
  • Automation Frameworks
  • Linux
  • Networking
  • Production Reliability
Soft Skills
  • Cross‑Team Collaboration
  • Technical Judgment
  • Problem‑Solving
Industry Keywords
  • Technical Strategy
  • Architectural Direction
  • Operational Standards
  • Production Readiness
  • Service Reliability
Tools & Technologies
  • Kubernetes
  • Cloud Platforms
  • Automation Tools
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff – Software Engineer, Infrastructure
Member of Technical Staff – Software Engineer, Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Product Engineer
Product Engineer

Jobtailor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

NVIDIA AI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Equity
Benefits
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Director Production Engineering
Director Production Engineering

Vista Applied Solutions Group Inc • Durham (NC)

On-site
USD 130,000 - 160,000
Cloud Architect, Solution Design Engineer
Cloud Architect, Solution Design Engineer

Jobtailor • California (MO)

On-site
USD 150,000 - 230,000
Senior Production Engineering Architect: Kubernetes & Automation
Senior Production Engineering Architect: Kubernetes & Automation

Jobtailor • California (MO)

On-site
USD 210,000 - 320,000
Distinguished Engineer, Production Engineering, Data Center Automation
Distinguished Engineer, Production Engineering, Data Center Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 320,000 - 489,000
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Distinguished Engineer, Production Engineering, Cluster Management
Distinguished Engineer, Production Engineering, Cluster Management

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 320,000 - 489,000