Senior Kubernetes Reliability Engineer - GPU Cloud

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 168,000 - 334,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI workloads. We seek Senior Reliability Engineers to help automate, safeguard, and scale production systems across cloud partners and on-prem environments.

We value strong Python/Go skills, deep Linux/Kubernetes expertise, and a solid grasp of SRE concepts. This role emphasizes observability, incident response, and cross-team collaboration in a fast-paced, GPU-focused setting.

Qualifications

  • 8+ years of experience building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar.
  • Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
  • Ability to troubleshoot distributed systems in production.
  • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
  • Clear communication and ability to work across teams.
  • BS/MS in Computer Science or equivalent experience.

Responsibilities

  • Build and operate automation for large-scale Kubernetes clusters across NVIDIA Cloud Partners (NCP) and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
  • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
  • Participate in on-call, incident response, debugging, and durable follow-up work.
  • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.

Skills

Python
Go
Kubernetes
Linux
SRE
Observability
GitOps
Incident management

Education

BS/MS in Computer Science or equivalent experience

Tools

OpenTelemetry
Prometheus
Grafana
ELK Stack
Lightstep
Splunk
Terraform
ArgoCD

Job description

NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI workloads. We seek Senior Reliability Engineers to help automate, safeguard, and scale production systems across cloud partners and on-prem environments.

We value strong Python/Go skills, deep Linux/Kubernetes expertise, and a solid grasp of SRE concepts. This role emphasizes observability, incident response, and cross-team collaboration in a fast-paced, GPU-focused setting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cloud Production Engineer (Remote)
Senior GPU Cloud Production Engineer (Remote)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Full health benefits
Retirement plan
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Cloud Orchestration Engineer
Senior Cloud Orchestration Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive Benefits
Senior Software Engineer, DGX Cloud Production - Equity
Senior Software Engineer, DGX Cloud Production - Equity

NVIDIA AI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Equity
Benefits
DGX Cloud Kubernetes Engineering Manager
DGX Cloud Kubernetes Engineering Manager

NVIDIA • Austin (TX)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Kubernetes Engineer for GPU AI Compute Platform
Senior Kubernetes Engineer for GPU AI Compute Platform

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
Principal Cloud Production Engineer: Kubernetes&Automation
Principal Cloud Production Engineer: Kubernetes&Automation

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Senior Software Engineer - DGX Cloud Production Engineering
Senior Software Engineer - DGX Cloud Production Engineering

NVIDIA AI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Equity
Benefits
GPU Cloud Platform Support Engineer: Kubernetes & Cloud Infra
GPU Cloud Platform Support Engineer: Kubernetes & Cloud Infra

NVIDIA • Santa Clara (CA)

On-site
USD 108,000 - 173,000
Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000