Senior SRE - GPU Cloud & Kubernetes, Equity Eligible

NVIDIA

Wyoming (OH)

On-site

USD 168,000 - 334,000

Full time

27 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA DGX Cloud is seeking Senior Reliability Engineers to build automation, tooling, and operational systems for GPU clusters. You will work on Kubernetes-based infrastructure, Day 2 operability, and GitOps across DGX Cloud environments.

The role requires 8+ years in production infrastructure, strong Python/Go skills, expertise in Linux/Kubernetes, and a solid grasp of SRE principles. Phase 2 stands out with GPU-infra and observability focus.

Qualifications

  • 8+ years of experience building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar.
  • Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
  • Ability to troubleshoot distributed systems in production.
  • Experience building and operating observability stacks (monitoring, logging, tracing).
  • BS/MS in Computer Science or equivalent experience.

Responsibilities

  • Build and operate automation for large-scale Kubernetes clusters across NVIDIA Cloud Partners and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
  • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting.
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
  • Participate in on-call, incident response, debugging, and durable follow-up work.
  • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.

Skills

Kubernetes
Python
Go
SRE principles
Incident management
Observability
Distributed systems
Linux

Education

BS/MS in Computer Science

Tools

OpenTelemetry
Prometheus
Grafana
ELK Stack
Splunk

Job description

NVIDIA DGX Cloud is seeking Senior Reliability Engineers to build automation, tooling, and operational systems for GPU clusters. You will work on Kubernetes-based infrastructure, Day 2 operability, and GitOps across DGX Cloud environments.

The role requires 8+ years in production infrastructure, strong Python/Go skills, expertise in Linux/Kubernetes, and a solid grasp of SRE principles. Phase 2 stands out with GPU-infra and observability focus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: GPU Cloud, Kubernetes & Automation
Senior SRE: GPU Cloud, Kubernetes & Automation

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior SRE for GPU Cloud: Kubernetes, Observability
Senior SRE for GPU Cloud: Kubernetes, Observability

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior GPU Cloud SRE — Kubernetes & Automation
Senior GPU Cloud SRE — Kubernetes & Automation

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Wyoming (OH)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Software Engineer, DGX Cloud Production - Equity
Senior Software Engineer, DGX Cloud Production - Equity

NVIDIA AI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Equity
Benefits
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

NVIDIA AI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Equity
Benefits
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000