Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe

Santa Clara (CA)

Hybrid

USD 248,000 - 397,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an experienced Site Reliability Engineer to lead the reliability roadmap for AI Platform Runtime and enterprise systems. You will architect highly available platforms, drive automation, and mentor senior engineers across multiple teams.

The role requires deep expertise in distributed systems, Kubernetes, and cloud platforms, with a focus on observability and incident-driven improvements. Hybrid work and equity opportunities are offered.

Qualifications

  • 15+ years in Site Reliability/Platform/Distributed Systems or related infra roles.
  • BS or MS in CS or related field with significant software development.
  • Proven experience leading large-scale engineering initiatives across multiple teams.
  • Deep expertise in distributed systems, networking, Linux, Kubernetes, and public clouds (AWS/Azure/GCP).
  • Proficiency in one or more programming languages (Python, Go, TypeScript, JavaScript, Java) with production automation experience.
  • Extensive experience with infrastructure-as-code and platform automation (Terraform, Crossplane, AWS CDK/CloudFormation).
  • Strong observability mindset with OpenTelemetry, metrics, logs, traces, and analytics.
  • Track record of improving availability, performance, and engineering productivity.

Responsibilities

  • Define and drive the long-term vision, architecture, and reliability roadmap for NVIDIA’s platforms.
  • Architect highly available, secure, and scalable distributed platforms powering AI products.
  • Lead design of AI agents, AI skills, and automation for operations and incident response.
  • Establish platform-wide reliability standards including SLOs, capacity models, and readiness requirements.
  • Identify systemic risks and lead cross-functional programs to improve availability and performance.
  • Advance observability across complex environments using OpenTelemetry and analytics.
  • Collaborate with senior leaders across orgs to inform architectural investment decisions.
  • Provide technical leadership during incidents and translate lessons into durable improvements.
  • Develop reference architectures and automation frameworks for broad adoption.
  • Mentor senior engineers and raise the technical bar across the company.

Skills

Distributed Systems
Kubernetes
Public Cloud (AWS/Azure/GCP)
Python
Go
TypeScript
JavaScript
Java
OpenTelemetry
Observability

Education

BS/MS in Computer Science or related field

Tools

Terraform
Crossplane
AWS CDK
AWS CloudFormation

Job description

NVIDIA is seeking an experienced Site Reliability Engineer to lead the reliability roadmap for AI Platform Runtime and enterprise systems. You will architect highly available platforms, drive automation, and mentor senior engineers across multiple teams.

The role requires deep expertise in distributed systems, Kubernetes, and cloud platforms, with a focus on observability and incident-driven improvements. Hybrid work and equity opportunities are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE - AI Platform Reliability Architect
Principal SRE - AI Platform Reliability Architect

NVIDIA AI • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Principal SRE: AI Platform Reliability Leader (Equity)
Principal SRE: AI Platform Reliability Leader (Equity)

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE — Scale AI Systems, Equity Eligible
Senior SRE — Scale AI Systems, Equity Eligible

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Staff SRE: AI Platform Runtime & Automation
Staff SRE: AI Platform Runtime & Automation

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000