Senior SRE: GPU Cloud, Kubernetes & Automation

NVIDIA

Santa Clara (CA)

On-site

USD 190,000 - 320,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA DGX Cloud is seeking a Senior Reliability Engineer to help build automation, tooling, and operational systems for GPU clusters in production. This role focuses on Kubernetes-based infrastructure, reliability, and Day 2 operability across DGX Cloud environments.

You will work on automation, observability, and incident management with a strong emphasis on Python/Go and Linux/Kubernetes expert knowledge.

Qualifications

  • 8+ years of experience building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar.
  • Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
  • Ability to troubleshoot distributed systems in production.
  • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

Responsibilities

  • Build and operate automation for large-scale Kubernetes clusters across NVIDIA Cloud Partners and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
  • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
  • Participate in on-call, incident response, debugging, and durable follow-up work.
  • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.

Skills

Python
Go
Linux
Kubernetes
Observability
SRE concepts
Incident management
Distributed systems

Education

BS/MS in Computer Science

Tools

OpenTelemetry
Prometheus
Grafana
ELK Stack
Lightstep
Splunk

Job description

NVIDIA DGX Cloud is seeking a Senior Reliability Engineer to help build automation, tooling, and operational systems for GPU clusters in production. This role focuses on Kubernetes-based infrastructure, reliability, and Day 2 operability across DGX Cloud environments.

You will work on automation, observability, and incident management with a strong emphasis on Python/Go and Linux/Kubernetes expert knowledge.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - GPU Cloud & Kubernetes, Equity Eligible
Senior SRE - GPU Cloud & Kubernetes, Equity Eligible

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE for GPU Cloud: Kubernetes, Observability
Senior SRE for GPU Cloud: Kubernetes, Observability

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior GPU Cloud SRE — Kubernetes & Automation
Senior GPU Cloud SRE — Kubernetes & Automation

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Software Engineer, DGX Cloud Production - Equity
Senior Software Engineer, DGX Cloud Production - Equity

NVIDIA AI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior SRE: Automate Reliability for AI/GPU Platform
Senior SRE: Automate Reliability for AI/GPU Platform

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
DGX Cloud Kubernetes Engineering Manager
DGX Cloud Kubernetes Engineering Manager

NVIDIA • Austin (TX)

On-site
USD 272,000 - 431,000
Equity
Benefits