Senior SRE — Scale AI Systems, Equity Eligible

NVIDIA AI

Santa Clara (CA)

On-site

USD 168,000 - 334,000

Full time

42 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA is seeking an experienced Site Reliability Engineer to lead the strategy and architecture of large-scale, enterprise-grade systems. You will design resilient distributed architectures powering NVIDIA’s AI-driven products and collaborate across Cloud, Platform, Security, and AI/ML teams to ensure high availability.

The role emphasizes automation, observability, and mentoring engineers to uphold reliability and performance across complex environments.

Qualifications

  • 10+ years in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
  • BS degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • Strong proficiency in Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.
  • Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform or CrossPlane.
  • Solid understanding of OpenTelemetry or other Observability implementation at scale.
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).
  • Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.

Responsibilities

  • Lead the technical strategy and roadmap for large-scale SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.
  • Design, and build resilient distributed systems that power NVIDIA’s next-generation AI-driven enterprise products and services.
  • Architect and develop AI Agents, AI Skills to accelerate platform operations.
  • Drive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.
  • Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.
  • Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.
  • Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.

Skills

Python
Typescript
JavaScript
Go
Automation

Education

BS in Computer Science or related field

Tools

Kubernetes
AWS CDK
CloudFormation
Terraform
CrossPlane

Job description

NVIDIA is seeking an experienced Site Reliability Engineer to lead the strategy and architecture of large-scale, enterprise-grade systems. You will design resilient distributed architectures powering NVIDIA’s AI-driven products and collaborate across Cloud, Platform, Security, and AI/ML teams to ensure high availability.

The role emphasizes automation, observability, and mentoring engineers to uphold reliability and performance across complex environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior SRE: Lead Reliable AI Infra with Equity
Senior SRE: Lead Reliable AI Infra with Equity

Nscale • Seattle (WA)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work environment
Principal SRE: AI Platform Reliability Leader (Equity)
Principal SRE: AI Platform Reliability Leader (Equity)

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Senior Staff SRE: Global Infra & Cloud Reliability
Senior Staff SRE: Global Infra & Cloud Reliability

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000