Senior AIOps SRE for AI Data Center Platform

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 148,000 - 276,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA Gruppe in Santa Clara is seeking an experienced engineer to build an AI Data Center AIOps platform. The ideal candidate will have a strong background in Kubernetes and automation, ensuring the reliability of GPU fleet management.

Key responsibilities include monitoring platform health, owning infrastructure deployments, and leading incident resolution. Candidates should possess 5+ years of experience in production systems and a degree in CS/CE. A competitive salary and generous benefits package are offered.

Qualifications

  • 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.
  • Proven ownership of reliability for an observability/AIOps platform.
  • Deep Kubernetes and containers experience for telemetry-heavy microservices.

Responsibilities

  • Continuously monitor platform health via dashboards, logs, and metrics.
  • Own Kubernetes deployments end-to-end including runbooks and validation.
  • Lead first-level incident triage and collect diagnostics.

Skills

Kubernetes
Python
Bash
Distributed systems
Automation

Education

BS/MS in CS/CE or equivalent experience

Tools

Terraform
Helm
Prometheus
Grafana

Job description

NVIDIA Gruppe in Santa Clara is seeking an experienced engineer to build an AI Data Center AIOps platform. The ideal candidate will have a strong background in Kubernetes and automation, ensuring the reliability of GPU fleet management.

Key responsibilities include monitoring platform health, owning infrastructure deployments, and leading incident resolution. Candidates should possess 5+ years of experience in production systems and a degree in CS/CE. A competitive salary and generous benefits package are offered.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE, AIOps Platform for GPU Data Centers
Senior SRE, AIOps Platform for GPU Data Centers

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Senior Site Reliability Engineer, AIOPs
Senior Site Reliability Engineer, AIOPs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Senior Site Reliability Engineer, AIOPs
Senior Site Reliability Engineer, AIOPs

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 148,000 - 276,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
AI Platform & SRE Engineering Leader
AI Platform & SRE Engineering Leader

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Inference Platform Engineer (Kubernetes & GPU)
Senior Inference Platform Engineer (Kubernetes & GPU)

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 241,500
Equity
Benefits
AI Platform & SRE Engineering Lead
AI Platform & SRE Engineering Lead

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000