Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA

Santa Clara (CA)

Hybrid

USD 168,000 - 334,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an experienced Site Reliability Engineer to lead the technical roadmap for large-scale, cross-functional SRE initiatives. You will design and build resilient distributed systems powering NVIDIA’s AI-driven enterprise products and collaborate with Cloud, Platform, Security, and AI/ML teams to ensure high availability and secure operations.

You will mentor engineers, drive automation, and advance observability practices, influencing across technical and organizational boundaries

Qualifications

  • BS degree in Computer Science or a related technical field, or equivalent experience.
  • 10+ years of experience in SRE/Platform Engineering/Cloud Architecture roles.
  • Strong programming skills in Python, TypeScript, JavaScript or Go with automation focus.
  • Experience with infrastructure-as-code: AWS CDK, CloudFormation, Terraform or CrossPlane.
  • Solid understanding of OpenTelemetry or other observability implementations at scale.
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS/Azure/GCP).
  • Excellent problem-solving, communication, and teamwork; ability to influence across teams.

Responsibilities

  • Lead technical strategy and roadmap for large-scale cross-functional SRE initiatives to improve reliability, scalability, and productivity.
  • Design and build resilient distributed systems powering NVIDIA's AI-driven products.
  • Architect and develop AI Agents and AI Skills to accelerate platform operations.
  • Drive automation and observability improvements using metrics to enhance performance and efficiency.
  • Collaborate across Cloud, Platform, Security, and AI/ML teams to enable high availability and secure operations.
  • Analyze and troubleshoot complex systems; champion best practices in incident management and postmortems.
  • Mentor engineers across teams to foster reliability-focused culture.

Skills

SRE
Platform engineering
Cloud architecture
Python
TypeScript
JavaScript
Go
IaC
OpenTelemetry
Kubernetes
Public cloud (AWS/Azure/GCP)
Networking

Education

BS in CS or related field

Tools

AWS CDK
AWS CloudFormation
Terraform
CrossPlane

Job description

NVIDIA is seeking an experienced Site Reliability Engineer to lead the technical roadmap for large-scale, cross-functional SRE initiatives. You will design and build resilient distributed systems powering NVIDIA’s AI-driven enterprise products and collaborate with Cloud, Platform, Security, and AI/ML teams to ensure high availability and secure operations.

You will mentor engineers, drive automation, and advance observability practices, influencing across technical and organizational boundaries

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal SRE - AI Platform Reliability & Automation
Principal SRE - AI Platform Reliability & Automation

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Hybrid work model
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits