Senior GPU Cloud SRE — Kubernetes & Automation

NVIDIA AI

Santa Clara (CA)

On-site

USD 168,000 - 334,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA DGX Cloud is seeking a Senior Reliability Engineer to design and operate automation for large GPU-based Kubernetes clusters across cloud partners and on-prem environments. You will shape observability, incident response, and production workflows to keep systems reliable and scalable.

You will implement SRE principles, define SLOs/SLIs, and collaborate across platform, storage, networking, and security teams. The role offers equity and a competitive total rewards package in Santa Clara, CA.

Qualifications

  • 8+ years of experience building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar.
  • Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
  • Ability to troubleshoot distributed systems in production.
  • Experience building and operating observability stacks (monitoring, logging, tracing).
  • Clear communication and ability to work across teams.
  • BS/MS in Computer Science or equivalent experience.

Responsibilities

  • Build and operate automation for large-scale Kubernetes clusters across NVIDIA Cloud Partners and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
  • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting.
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
  • Participate in on-call, incident response, debugging, and durable follow-up work.
  • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.

Skills

Python
Go
Linux
Kubernetes
SRE
Observability
OpenTelemetry
Prometheus

Education

BS/MS in Computer Science or equivalent

Tools

OpenTelemetry
Prometheus
Grafana
ELK Stack
Lightstep
Splunk

Job description

NVIDIA DGX Cloud is seeking a Senior Reliability Engineer to design and operate automation for large GPU-based Kubernetes clusters across cloud partners and on-prem environments. You will shape observability, incident response, and production workflows to keep systems reliable and scalable.

You will implement SRE principles, define SLOs/SLIs, and collaborate across platform, storage, networking, and security teams. The role offers equity and a competitive total rewards package in Santa Clara, CA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE for GPU Cloud: Kubernetes, Observability
Senior SRE for GPU Cloud: Kubernetes, Observability

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: GPU Cloud, Kubernetes & Automation
Senior SRE: GPU Cloud, Kubernetes & Automation

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior SRE - GPU Cloud & Kubernetes, Equity Eligible
Senior SRE - GPU Cloud & Kubernetes, Equity Eligible

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: Automate Reliability for AI/GPU Platform
Senior SRE: Automate Reliability for AI/GPU Platform

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity options
Comprehensive benefits
Principal Cloud Production Engineer: Kubernetes&Automation
Principal Cloud Production Engineer: Kubernetes&Automation

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250