Senior Production Engineer - AI Cloud & DGX Ops

NVIDIA Corporation

Zürich

Hybrid

CHF 140,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA Corporation is seeking a Senior Production Engineer to build software and automation for DGX Cloud, ensuring reliability, scalability, and safe operation across model endpoints and compute infrastructure.

You will work on large-scale distributed systems spanning Kubernetes clusters across public clouds and on-premises, with a focus on health validation, rollouts, and observability.

Qualifications

  • 8+ years of experience building or operating production services and large-scale distributed systems.
  • Strong programming skills in Python or Go with experience building production tooling.
  • Experience with infrastructure as code and GitOps for repeatable deployments.
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, networking fundamentals.
  • Experience instrumenting services with metrics, logs, and traces to improve reliability.

Responsibilities

  • Build and operate production software, automation, and tooling for DGX Cloud control plane services and model deployments.
  • Improve availability, routing, capacity management, and observability of inference and agentic workloads.
  • Define and instrument SLIs/SLOs for inference and control plane services; monitor error budgets.
  • Collaborate with model, platform, storage, networking, security, and GPU infra teams to operate services safely at scale.
  • Participate in on-call and incident response, turning recurring issues into automation.

Skills

Python
Go
Infrastructure as Code
GitOps
Linux
Kubernetes
Distributed systems

Education

BS/MS in Computer Science

Tools

Terraform
Argo CD
Docker

Job description

NVIDIA Corporation is seeking a Senior Production Engineer to build software and automation for DGX Cloud, ensuring reliability, scalability, and safe operation across model endpoints and compute infrastructure.

You will work on large-scale distributed systems spanning Kubernetes clusters across public clouds and on-premises, with a focus on health validation, rollouts, and observability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Production Engineer - AI Cloud Reliability
Senior Production Engineer - AI Cloud Reliability

NVIDIA • Zürich

On-site
CHF 140,000 - 180,000
Senior Production Engineer - DGX Cloud
Senior Production Engineer - DGX Cloud

NVIDIA • Zürich

On-site
CHF 140,000 - 180,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Switzerland

On-site
CHF 150,000 - 210,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Corporation • Zürich

On-site
CHF 180,000 - 230,000
Senior SRE: AI Cloud Infra & Kubernetes
Senior SRE: AI Cloud Infra & Kubernetes

NVIDIA Gruppe • Zürich

On-site
CHF 180,000 - 230,000
Remote Senior Performance Engineer - AI & HPC Systems
Remote Senior Performance Engineer - AI & HPC Systems

NVIDIA Corporation • Zürich

On-site
CHF 140,000 - 210,000
Senior HPC and AI Network Software Architect
Senior HPC and AI Network Software Architect

NVIDIA Corporation • Zürich

On-site
CHF 180,000 - 240,000
Senior HPC and AI Network Software Architect
Senior HPC and AI Network Software Architect

NVIDIA • Zürich

On-site
CHF 180,000 - 240,000
Senior HPC AI Network Architect: Scalable Infra
Senior HPC AI Network Architect: Scalable Infra

NVIDIA Corporation • Zürich

On-site
CHF 180,000 - 240,000
Senior HPC & AI Network Architect for Scalable AI Infra
Senior HPC & AI Network Architect for Scalable AI Infra

NVIDIA • Zürich

On-site
CHF 180,000 - 240,000