Lead AI Platform Reliability Architect

NVIDIA

Santa Clara (CA)

Hybrid

USD 248,000 - 397,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is looking for a Principal SRE to shape the technical direction of the AI Platform Runtime and lead reliability initiatives across the organization. You will design highly available, scalable platforms and develop AI-driven automation that accelerates operations and incident response.

You will establish SLOs, capacity models, and observability practices using OpenTelemetry, while mentoring senior engineers and collaborating with Cloud, Platform, Security, Networking, and AI/ML teams to

Qualifications

  • 15+ years of experience in SRE/Platform engineering or related roles.
  • BS or MS degree in Computer Science or related field.
  • Experience setting technical strategy across multiple teams or organizations.
  • Deep expertise in distributed systems, Linux, Kubernetes, and cloud platforms (AWS/Azure/GCP).
  • Strong programming skills in Python/Go/TypeScript/JavaScript/Java with production automation experience.
  • Experience with infrastructure-as-code and platform automation tools (Terraform, Crossplane, AWS CDK/CloudFormation).
  • Deep understanding of observability with OpenTelemetry, metrics, logging, tracing, and analytics.
  • Proven track record delivering improvements in availability, performance, or productivity.
  • Strong communication and collaboration skills to influence stakeholders across boundaries.

Responsibilities

  • Define and drive the long-term technical vision, architecture, and roadmap for reliability across NVIDIA’s AI Platform Runtime and related systems.
  • Architect highly available, resilient, secure, and scalable distributed platforms for AI-driven products and services.
  • Lead the design and development of AI agents and automation to accelerate platform operations and incident response.
  • Establish platform-wide reliability standards, including SLOs, error budgets, and capacity models.
  • Identify systemic risks and lead cross-functional programs to improve availability, scalability, and developer productivity.
  • Advance observability through OpenTelemetry, metrics, logs, traces, profiling, and analytics.
  • Collaborate with Cloud, Platform, Security, Networking, and AI/ML teams on architectural decisions.
  • Provide technical leadership during incidents and translate lessons into durable improvements.
  • Develop reference architectures and automation frameworks adopted across engineering groups.
  • Mentor senior engineers and raise the technical bar across the company.

Skills

Distributed Systems
Kubernetes
Cloud Architecture
Programming: Python/Go/TypeScript/Java
OpenTelemetry
Reliability Engineering

Education

BS/MS in Computer Science or related field

Tools

Terraform
AWS CDK
CloudFormation
Crossplane

Job description

NVIDIA is looking for a Principal SRE to shape the technical direction of the AI Platform Runtime and lead reliability initiatives across the organization. You will design highly available, scalable platforms and develop AI-driven automation that accelerates operations and incident response.

You will establish SLOs, capacity models, and observability practices using OpenTelemetry, while mentoring senior engineers and collaborating with Cloud, Platform, Security, Networking, and AI/ML teams to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE - AI Platform Reliability & Automation
Principal SRE - AI Platform Reliability & Automation

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Senior Platform Architect – AI/ML Infrastructure
Senior Platform Architect – AI/ML Infrastructure

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Principal Infra Automation Engineer - AI Platform & Equity
Principal Infra Automation Engineer - AI Platform & Equity

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 248,000 - 391,000
Equity
Benefits
Principal Architect for Agentic AI Platforms & Foundations
Principal Architect for Agentic AI Platforms & Foundations

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits