Principal SRE - AI Platform Reliability & Automation

NVIDIA Corporation

Santa Clara (CA)

Hybrid

USD 248,000 - 397,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits
Hybrid work model

Job summary

NVIDIA Corporation is seeking a Principal SRE to steer the reliability strategy for its AI Platform Runtime. You will architect highly available, scalable distributed platforms and lead the development of AI agents and automation to improve platform operations and incident response.

You will collaborate across Cloud, Platform, Security, Networking, and AI/ML teams to set standards and drive durable improvements. Equity and benefits accompany a base salary and hybrid work options.

Qualifications

  • 15+ years in SRE, platform engineering, distributed systems or related roles.
  • Degree in CS or equivalent software development experience.
  • Proven strategy leadership across multiple teams.
  • Deep expertise in distributed systems, networking, Linux, Kubernetes, and public clouds (AWS/Azure/GCP).
  • Strong programming skills: Python, Go, TypeScript/JavaScript, or Java with automation focus.
  • IaC and platform automation: Terraform, Crossplane, AWS CDK/CloudFormation.
  • Observability at scale with OpenTelemetry, metrics, logs, traces, analytics.
  • Reliability engineering practices: SLOs, budgets, capacity, incident management, postmortems.
  • Proven impact on availability, performance, or productivity in complex environments.
  • Excellent judgment, communication, and collaboration to influence stakeholders.

Responsibilities

  • Define long-term reliability vision, architecture, and roadmaps for NVIDIA's AI Platform Runtime.
  • Design highly available, secure, scalable distributed platforms for AI-driven products.
  • Lead AI agents, AI skills, and automation to speed platform operations and incident response.
  • Establish platform-wide reliability standards, SLOs, error budgets, capacity models.
  • Identify systemic risks and drive cross-functional improvement programs.
  • Advance observability with OpenTelemetry, metrics, logs, tracing, and analytics.
  • Partner with Cloud, Platform, Security, Networking, and AI/ML teams for architectural decisions.
  • Provide technical leadership during incidents and translate learnings into durable changes.
  • Develop reference architectures and automation frameworks for multiple teams.
  • Mentor engineers and raise the technical bar across the company.

Skills

SRE leadership
Distributed systems
Cloud platforms
Kubernetes
Networking
Python
Go
TypeScript
Java
Observability
OpenTelemetry

Education

BS or MS in Computer Science

Tools

Terraform
Crossplane
AWS CDK
AWS CloudFormation
OpenTelemetry

Job description

NVIDIA Corporation is seeking a Principal SRE to steer the reliability strategy for its AI Platform Runtime. You will architect highly available, scalable distributed platforms and lead the development of AI agents and automation to improve platform operations and incident response.

You will collaborate across Cloud, Platform, Security, Networking, and AI/ML teams to set standards and drive durable improvements. Equity and benefits accompany a base salary and hybrid work options.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Principal AI Platform Engineer – Agentic Apps & Foundations
Principal AI Platform Engineer – Agentic Apps & Foundations

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Principal SRE: AI Inference Reliability & Scale
Principal SRE: AI Inference Reliability & Scale

Cerebras • Sunnyvale (CA)

On-site
USD 260,000 - 380,000
Principal Infra Automation Engineer - AI Platform & Equity
Principal Infra Automation Engineer - AI Platform & Equity

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 248,000 - 391,000
Equity
Benefits