AI Platform & SRE Engineering Lead

Jobs in JS

Santa Clara (CA)

On-site

USD 208,000 - 334,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NVIDIA in Santa Clara, CA is seeking an Engineering Manager to lead a team of SREs and platform engineers building and operating resilient AI platform capabilities at enterprise scale.

You will shape strategy, mentor engineers, and partner across Cloud, Security, Networking, and AI/ML teams to improve reliability and developer productivity. The role emphasizes blameless postmortems and fostering an inclusive, high-performing culture.

Qualifications

  • 10+ years in SRE or platform engineering.
  • 3+ years leading or managing engineering teams.
  • Strong distributed systems, Linux, networking, Kubernetes experience.
  • Experience with public clouds (AWS, Azure, GCP).
  • Solid observability with metrics, logs, tracing.
  • Proven ability to deliver roadmaps and measurable outcomes.
  • Excellent communication and cross-functional collaboration.
  • Experience with incident response, disaster recovery, blameless postmortems.
  • Experience building AI-related platform capabilities is a plus.

Responsibilities

  • Lead and grow a team of SRE, platform, and software engineers for AI Platform Runtime.
  • Define the team’s technical strategy, priorities, and roadmap aligned with product and business objectives.
  • Guide design and delivery of highly available, scalable distributed systems.
  • Drive development of AI agents, AI skills, and automation for platform operations.
  • Establish reliability goals and operational practices with SLIs/SLOs, error budgets, capacity models, health metrics.
  • Improve velocity and developer experience via self-service platforms and IaC.
  • Coordinate with product, architecture, and leaders across Cloud, Security, Networking, Platform, AI/ML.
  • Balance feature delivery, platform investment, and technical debt.
  • Lead incidents and ensure blameless postmortems lead to ownership and fixes.
  • Recruit and develop engineers, providing coaching and career growth.

Skills

SRE leadership
Distributed systems
Linux
Kubernetes
Public cloud
Observability
Communication
Hiring
Coaching
Roadmap planning

Tools

Python
Go
TypeScript
JavaScript
Java
OpenTelemetry

Job description

NVIDIA in Santa Clara, CA is seeking an Engineering Manager to lead a team of SREs and platform engineers building and operating resilient AI platform capabilities at enterprise scale.

You will shape strategy, mentor engineers, and partner across Cloud, Security, Networking, and AI/ML teams to improve reliability and developer productivity. The role emphasizes blameless postmortems and fostering an inclusive, high-performing culture.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform & SRE Engineering Leader
AI Platform & SRE Engineering Leader

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
AI Platform & SRE Engineering Manager
AI Platform & SRE Engineering Manager

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Engineering Manager - AI Platform & SRE
Engineering Manager - AI Platform & SRE

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Engineering Manager – AI Platform & SRE
Engineering Manager – AI Platform & SRE

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Engineering Manager – AI Platform & SRE
Engineering Manager – AI Platform & SRE

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits