Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a Senior SRE to join the Compute Farm team, responsible for end-to-end SRE solutions across a multi-cloud hybrid environment. You will drive design, automation, and reliability for critical services while collaborating with global teams.

Ideal candidates bring 5+ years of HPC experience, strong coding skills, and a track record mentoring engineers. The role emphasizes on-call participation, incident reviews, and data-driven incident resolution.

Qualifications

  • B.S. degree in Computer Science or related field, or equivalent experience.
  • 5+ years building and supporting critical services.
  • Experience supporting large HPC clusters with Slurm/LSF/Kubernetes.
  • Proficiency with CI/CD and Infrastructure as Code.
  • Experience with auto-healing, fleet management and observability.
  • Proficient in monitoring, metrics, container management, log collection.
  • 5+ years coding/scripting in Python, Go, Perl, or Ruby.
  • Mentored engineers and influenced technical direction.
  • Creative problem solving and strong communication skills.

Responsibilities

  • Own SRE solutions end-to-end from design to operation and continuous improvement.
  • Use IaC and config management to automate provisioning everywhere.
  • Deliver solutions in a globally distributed, multi-cloud hybrid environment (On-prem, AWS, GCP, OCI).
  • Design for failure with redundancy, failure domains, and change control.
  • Ensure uptime and QoS for internal customers through operational excellence.
  • Conduct capacity management and planning to meet ongoing needs.
  • Detect performance issues and propose solutions to maintain service quality.
  • Collaborate with teams to ensure seamless project completion.
  • Participate in on-call, incident reviews, and produce RCA reports.

Skills

SRE experience
HPC clusters
CI/CD
IaC
Observability
Monitoring
Container management
Logging/metrics
Python
Go
Mentoring
Communication

Education

B.S. degree in Computer Science or related field

Tools

Kubernetes
CI/CD pipelines
Log collection tools

Job description

NVIDIA is seeking a Senior SRE to join the Compute Farm team, responsible for end-to-end SRE solutions across a multi-cloud hybrid environment. You will drive design, automation, and reliability for critical services while collaborating with global teams.

Ideal candidates bring 5+ years of HPC experience, strong coding skills, and a track record mentoring engineers. The role emphasizes on-call participation, incident reviews, and data-driven incident resolution.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior SRE — Global HPC/AI Infra with Equity
Senior SRE — Global HPC/AI Infra with Equity

Socket.dev • North Carolina

On-site
USD 152,000 - 288,000
Senior Staff SRE: Global On-Prem & Cloud Infra Lead
Senior Staff SRE: Global On-Prem & Cloud Infra Lead

NVIDIA • United States

On-site
USD 200,000 - 322,000
Senior HPC Storage SRE: On-Prem & Cloud Automation
Senior HPC Storage SRE: On-Prem & Cloud Automation

JobCubby • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: Global Infra & Automation (Equity)
Senior SRE: Global Infra & Automation (Equity)

Socket.dev • California (MO)

Hybrid
USD 200,000 - 322,000
Equity
Benefits
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA Corporation • Durham (CA), Northern (KY)

On-site
USD 152,000 - 288,000
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

Socket.dev • North Carolina

On-site
USD 152,000 - 288,000
Senior SRE - Scalable Infra & Reliability (Equity)
Senior SRE - Scalable Infra & Reliability (Equity)

NVIDIA • Durham (NC)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1