Principal SRE - AI Platform Reliability Architect

NVIDIA AI

Santa Clara (CA)

On-site

USD 248,000 - 397,000

Full time

12 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA is seeking a Principal SRE to shape the technical direction of its AI Platform Runtime and lead reliability initiatives across multiple teams. You will design highly available distributed platforms and advance observability to support AI services at scale.

You will mentor engineers, drive improvements in availability, performance, and developer productivity, and collaborate with Cloud, Platform, Security, Networking, and AI/ML groups.

Qualifications

  • 15+ years in SRE, platform engineering, or related infrastructure roles.
  • BS or MS in CS or a related technical field, or equivalent experience.
  • Proven track record leading large-scale engineering initiatives across teams.
  • Deep expertise in distributed systems, networking, Linux, Kubernetes, and cloud platforms (AWS/Azure/GCP).
  • Strong programming skills in Python/Go/TypeScript/Java; experience with automation.
  • Experience with IaC and platform automation tools (Terraform, Crossplane, AWS CDK/CF).
  • Expertise in observability including OpenTelemetry, metrics, logs, traces.
  • Proven impact on availability, performance, or productivity.

Responsibilities

  • Define long-term vision, architecture, and roadmap for reliability.
  • Architect highly available, resilient, secure, and scalable platforms.
  • Develop AI agents and automation to accelerate operations.
  • Establish platform-wide reliability standards and SLOs.
  • Lead cross-functional programs to improve availability and performance.
  • Improve observability across complex environments with OpenTelemetry.
  • Partner with Cloud, Platform, Security, Networking, AI/ML teams.
  • Provide technical leadership during incidents.
  • Develop reference architectures and automation frameworks.
  • Mentor senior engineers and raise technical bar.

Skills

SRE
Distributed Systems
Cloud Architecture
Kubernetes
OpenTelemetry
Programming Languages

Education

BS/MS in CS or related field

Tools

Terraform
Crossplane
AWS CDK
CloudFormation

Job description

NVIDIA is seeking a Principal SRE to shape the technical direction of its AI Platform Runtime and lead reliability initiatives across multiple teams. You will design highly available distributed platforms and advance observability to support AI services at scale.

You will mentor engineers, drive improvements in availability, performance, and developer productivity, and collaborate with Cloud, Platform, Security, Networking, and AI/ML groups.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE - AI Platform Reliability & Automation
Principal SRE - AI Platform Reliability & Automation

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior SRE: Lead Reliable AI Platform & Mentor the Team
Senior SRE: Lead Reliable AI Platform & Mentor the Team

Nscale • Houston (TX)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work culture
Principal SRE: AI Inference Reliability & Scale
Principal SRE: AI Inference Reliability & Scale

Cerebras • Sunnyvale (CA)

On-site
USD 260,000 - 380,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model