Turn this role into an interview — a resume and cover letter built around what this employer wants.
NVIDIA is seeking a Distinguished Engineer to lead cluster operations for DGX Cloud capacity across on-prem and cloud environments.
The role blends software engineering, systems knowledge, and production discipline to build reliable, scalable DGX Cloud platforms and workflows that serve researchers and customers.
You will define technical strategy, set operating standards, and drive cross-team delivery to improve production readiness and overall platform availability at scale.
NVIDIA is looking for a Distinguished Engineer to act as a senior technical leader in the Production Engineering group, enthusiastic about cluster operations involving DGX Cloud GPU capacity.
At NVIDIA, Production Engineering is responsible for ensuring large-scale production systems are reliable, straightforward to lead, and increasingly automated across NVIDIA's DGX Cloud resources. We combine software engineering, systems engineering, and extensive production knowledge to build platforms, workflows, and operational frameworks that sustain GPU infrastructure health, scalability, and availability for researchers and customers.
This role centers on the operational framework supporting DGX Cloud environments spanning on-premises, major cloud providers, and NVIDIA Cloud Partner locations. The responsibilities include engineering integrations to guarantee DGX Cloud capacity is fully operational in production. This involves Kubernetes service management, ensuring vendor and equipment availability, on-prem infrastructure operations, release and runtime preparation, and service reliability coordination. The workflows tie these elements into a cohesive production system.
This hands-on Distinguished Engineer role calls for a deeply technical leader to build the architectural direction for cluster operations in DGX Cloud. The ideal candidate will blend software engineering expertise, system knowledge, and production insight to define technical strategy, set operating standards, direct the evolution of the production model, and drive delivery of cross-organizational capabilities. These capabilities ensure that the DGX Cloud resources remain usable, maintainable, and improve continuously at scale. The position demands both deep invention and implementation skills and the ability to lead by influence across several teams and critical production results.
Define the long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments
Define the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability throughout DGX Cloud resources
Guide the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination
Make and guide high-impact technical decisions that resolve how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production
Develop robust workflows, interfaces, and engineering collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability domains