AI Platform & SRE Engineering Manager

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 208,000 - 334,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA Corporation in Santa Clara, CA, is seeking an Engineering Manager for Site Reliability Engineering to lead a team responsible for NVIDIA's AI Platform Runtime and related production services. You will combine people leadership with technical judgment to deliver reliable, scalable systems at enterprise scale.

You will partner with Cloud, Platform, Security, and AI/ML organizations to advance platform reliability, developer productivity, and automation, while mentoring engineers and

Qualifications

  • 10+ years in SRE, Platform Eng, or related field with 3+ years leading engineering teams.
  • Distributed systems, Linux, networking, Kubernetes, and public cloud expertise.
  • Leading teams that build production software and automation using Python, Go, TypeScript/JavaScript, or Java.
  • Strong observability at scale: OpenTelemetry, metrics, logs, tracing, analytics.
  • Experience applying SRE practices: SLOs, incident mgmt, capacity, blameless postmortems.

Responsibilities

  • Lead and grow a team of SREs, platform, and software engineers for NVIDIA's AI Platform Runtime.
  • Define team strategy, priorities, and roadmap aligned with product, platform, and business goals.
  • Design highly available, scalable, secure, and resilient distributed systems.
  • Drive AI agents, AI skills, and automation for platform operations and incident response.
  • Improve velocity via self-service platforms, IaC, and standardized delivery.
  • Coordinate with Cloud, Platform, Security, Networking, and AI/ML leaders to drive cross-functional initiatives.

Skills

SRE leadership
Distributed systems
Linux
Networking
Kubernetes
Public cloud (AWS/Azure/GCP)
Programming languages (Python, Go, TS/
Observability
SRE practices (SLOs, error budgets)
Roadmapping & cross-team collaboration

Job description

NVIDIA Corporation in Santa Clara, CA, is seeking an Engineering Manager for Site Reliability Engineering to lead a team responsible for NVIDIA's AI Platform Runtime and related production services. You will combine people leadership with technical judgment to deliver reliable, scalable systems at enterprise scale.

You will partner with Cloud, Platform, Security, and AI/ML organizations to advance platform reliability, developer productivity, and automation, while mentoring engineers and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform & SRE Engineering Leader
AI Platform & SRE Engineering Leader

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
AI Platform & SRE Engineering Lead
AI Platform & SRE Engineering Lead

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Engineering Manager - AI Platform & SRE
Engineering Manager - AI Platform & SRE

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Engineering Manager – AI Platform & SRE
Engineering Manager – AI Platform & SRE

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Engineering Manager – AI Platform & SRE
Engineering Manager – AI Platform & SRE

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 248,000 - 397,000