AI Platform & SRE Engineering Leader

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 208,000 - 334,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA Corporation in Santa Clara, CA is seeking an Engineering Manager to lead a team of SRE, platform, and software engineers responsible for the AI Platform Runtime and related production services. You will set the team's strategy, roadmaps, and drive highly available, scalable, and secure systems for enterprise AI products.

You will recruit and develop engineers, partner with Cloud, Security, Networking, and AI/ML teams, improve developer productivity through automation, and guide incident

Qualifications

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, Software Engineering, Cloud Infrastructure, or a related technical field.
  • 3+ years managing or formally leading engineering teams responsible for complex production systems.
  • Technical foundation in distributed systems, Linux, networking, Kubernetes, and public cloud platforms such as AWS, Azure, or GCP.

Responsibilities

  • Lead, develop, and grow a team of SRE, platform, and software engineers responsible for NVIDIA's AI Platform Runtime and related production services.
  • Define the team's technical strategy, priorities, and roadmap in alignment with broader product, platform, and business objectives.
  • Guide the design and delivery of highly available, scalable, secure, and resilient distributed systems that support enterprise AI agent products.
  • Drive the development of AI agents, AI skills, and intelligent automation for platform operations, incident response, troubleshooting, and remediation.
  • Establish measurable reliability goals and effective operational practices using service-level indicators, service-level objectives, error budgets, capacity models, operational health metrics, and production readiness reviews.

Skills

Distributed systems
Linux
Networking
Kubernetes
Public cloud
Observability
Python/Go/TS/JS/Java
OpenTelemetry
SRE practices
Roadmapping/strategy

Tools

AWS
Azure
GCP
CI/CD

Job description

NVIDIA Corporation in Santa Clara, CA is seeking an Engineering Manager to lead a team of SRE, platform, and software engineers responsible for the AI Platform Runtime and related production services. You will set the team's strategy, roadmaps, and drive highly available, scalable, and secure systems for enterprise AI products.

You will recruit and develop engineers, partner with Cloud, Security, Networking, and AI/ML teams, improve developer productivity through automation, and guide incident

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform & SRE Engineering Lead
AI Platform & SRE Engineering Lead

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
AI Platform & SRE Engineering Manager
AI Platform & SRE Engineering Manager

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Engineering Manager – AI Platform & SRE
Engineering Manager – AI Platform & SRE

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Engineering Manager – AI Platform & SRE
Engineering Manager – AI Platform & SRE

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Engineering Manager - AI Platform & SRE
Engineering Manager - AI Platform & SRE

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits