Senior SRE: Global HPC & Cloud Reliability

NVIDIA AI

Durham (NC)

On-site

USD 184,000 - 288,000

Full time

1 hour ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a Senior SRE to join the Compute Farm team to help build the next generation of our global services platform. You will own SRE solutions end‑to‑end, from design and implementation to operation and continuous improvement, ensuring integration with HPC schedulers, storage, and network fabrics.

You’ll deliver in a globally distributed, multi‑cloud hybrid environment—On‑prem, AWS, GCP, and OCI— while designing for failure, redundancy, and strong QoS across internal customers.

Qualifications

  • BS in Computer Science or related technical field or equivalent experience.
  • 5+ years of professional experience building and supporting critical services.
  • Experience with large-scale HPC clusters and schedulers (Slurm/LSF/Kubernetes).
  • Strong CI/CD and Infrastructure as Code practices.
  • Proven ability to design scalable, observable platforms and auto-healing.

Responsibilities

  • Own SRE solutions end-to-end from design to operation and improvement.
  • Automate provisioning across on‑prem and multi‑cloud environments.
  • Deliver reliable services with high uptime and QoS for internal customers.
  • Lead on-call and incident reviews; produce high‑quality RCA reports.
  • Mentor engineers and influence architectural decisions.

Skills

SRE
HPC
Kubernetes
Python
Go
IaC
CI/CD
Observability
Incident response
Mentoring

Education

BS in Computer Science or related field

Tools

Slurm
LSF
Kubernetes

Job description

NVIDIA is seeking a Senior SRE to join the Compute Farm team to help build the next generation of our global services platform. You will own SRE solutions end‑to‑end, from design and implementation to operation and continuous improvement, ensuring integration with HPC schedulers, storage, and network fabrics.

You’ll deliver in a globally distributed, multi‑cloud hybrid environment—On‑prem, AWS, GCP, and OCI— while designing for failure, redundancy, and strong QoS across internal customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior SRE — Global HPC/AI Infra with Equity
Senior SRE — Global HPC/AI Infra with Equity

Socket.dev • North Carolina

On-site
USD 152,000 - 288,000
Senior HPC Storage SRE — On-Prem & Cloud, Equity
Senior HPC Storage SRE — On-Prem & Cloud, Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits package
Senior HPC Storage SRE — On-Prem + Cloud, Equity
Senior HPC Storage SRE — On-Prem + Cloud, Equity

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Equity
Benefits package
Senior Staff SRE: Global On-Prem & Cloud Infra Lead
Senior Staff SRE: Global On-Prem & Cloud Infra Lead

NVIDIA • United States

On-site
USD 200,000 - 322,000
Senior Staff SRE: Global Infra & Cloud Reliability
Senior Staff SRE: Global Infra & Cloud Reliability

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior HPC Storage SRE: On-Prem & Cloud Automation
Senior HPC Storage SRE: On-Prem & Cloud Automation

JobCubby • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA Corporation • Durham (CA), Northern (KY)

On-site
USD 152,000 - 288,000
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA AI • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits package
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

Socket.dev • North Carolina

On-site
USD 152,000 - 288,000