Senior SRE for GPU Cloud: Kubernetes, Observability

NVIDIA

California (MO)

On-site

USD 168,000 - 334,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI workloads. We seek a Senior Reliability Engineer to drive automation, tooling, and production systems across Kubernetes-based clusters and DGX environments.

You will develop observability stacks, define SLOs/SLIs, and partner with cross-functional teams to ensure production readiness. A strong background in SRE, distributed systems, and cloud-native ops is required.

Qualifications

  • 8+ years building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar.
  • Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
  • Ability to troubleshoot distributed systems in production.
  • Experience building and operating observability stacks (monitoring, logging, tracing).
  • Clear communication and ability to work across teams.
  • BS/MS in Computer Science or equivalent experience.

Responsibilities

  • Build and operate automation for large-scale Kubernetes clusters across NVIDIA Cloud Partners (NCP) and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
  • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
  • Participate in on-call, incident response, debugging, and durable follow-up work.
  • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.

Skills

Python
Go
Linux
Kubernetes
CloudInfrastructure
SREPrinciples
DistributedSystems
Observability

Education

BS/MS in CS or equivalent

Tools

OpenTelemetry
Prometheus
Grafana
ELK Stack
Lightstep
Splunk
Terraform
ArgoCD

Job description

NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI workloads. We seek a Senior Reliability Engineer to drive automation, tooling, and production systems across Kubernetes-based clusters and DGX environments.

You will develop observability stacks, define SLOs/SLIs, and partner with cross-functional teams to ensure production readiness. A strong background in SRE, distributed systems, and cloud-native ops is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: GPU Cloud, Kubernetes & Automation
Senior SRE: GPU Cloud, Kubernetes & Automation

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior GPU Cloud SRE — Kubernetes & Automation
Senior GPU Cloud SRE — Kubernetes & Automation

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior SRE - GPU Cloud & Kubernetes, Equity Eligible
Senior SRE - GPU Cloud & Kubernetes, Equity Eligible

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: Automate Reliability for AI/GPU Platform
Senior SRE: Automate Reliability for AI/GPU Platform

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior SRE - Customer-Facing GPU Cloud Reliability
Senior SRE - Customer-Facing GPU Cloud Reliability

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits