Senior SRE, AIOps Platform for GPU Data Centers

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 148,000 - 276,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA Corporation in Santa Clara, CA is hiring a DevOps Engineer to operate our AI Data Center telemetry platform. You’ll own reliability, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs/dashboards used by operators.

Expect to lead Kubernetes deployments end-to-end, build runbooks, and partner with Software and Systems Engineering to translate platform signals into actionable, trustworthy alerts and automation.

Qualifications

  • BS/MS in CS/CE (or equivalent experience) required.
  • 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.

Responsibilities

  • Continuously monitor platform health via dashboards/logs/metrics.
  • Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation).
  • Lead first-level incident triage and hand off actionable findings to engineering.
  • Build and maintain runbooks/SOPs/checklists; drive automation improvement.
  • Manage deployment infrastructure and packaging (Helm + Terraform) for scalable environments.
  • Contribute in adjacent areas to grow the team and knowledge.

Skills

Kubernetes expertise
Python scripting
Automation tooling
Observability & monitoring
CI/CD
Runbooks & SOPs
Distributed systems

Education

BS/MS in CS/CE or equivalent experience

Tools

Kubernetes
Terraform
Helm
Prometheus
Grafana
Kafka/Pulsar
ClickHouse/Elastic/TSDBs

Job description

NVIDIA Corporation in Santa Clara, CA is hiring a DevOps Engineer to operate our AI Data Center telemetry platform. You’ll own reliability, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs/dashboards used by operators.

Expect to lead Kubernetes deployments end-to-end, build runbooks, and partner with Software and Systems Engineering to translate platform signals into actionable, trustworthy alerts and automation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AIOps SRE for AI Data Center Platform
Senior AIOps SRE for AI Data Center Platform

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior AI Infrastructure Engineer Observability & Automation
Senior AI Infrastructure Engineer Observability & Automation

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 356,000
Equity
Benefits
Senior Site Reliability Engineer, AIOPs
Senior Site Reliability Engineer, AIOPs

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 148,000 - 276,000
AI Platform & SRE Engineering Lead
AI Platform & SRE Engineering Lead

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
AI Platform & SRE Engineering Leader
AI Platform & SRE Engineering Leader

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Senior Site Reliability Engineer, AIOPs
Senior Site Reliability Engineer, AIOPs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Senior AI/HPC Telemetry & Observability Engineer
Senior AI/HPC Telemetry & Observability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Senior AIOps & Observability Platform Architect
Senior AIOps & Observability Platform Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior SRE & Automation Architect — GPU Cloud, Remote
Senior SRE & Automation Architect — GPU Cloud, Remote

Bitdeer Technologies Group • San Jose (CA)

Remote
USD 180,000 - 260,000