Senior HPC SRE: Global Multi-Cloud Reliability

NVIDIA

Austin (TX)

On-site

USD 152,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Comprehensive benefits package

Job summary

A leading technology company in Austin, Texas is seeking a Senior SRE to join their Compute Farm team. In this role, you will be responsible for owning SRE solutions and ensuring the highest uptime for missions-critical services. The ideal candidate should have at least 5 years of experience in supporting large-scale HPC clusters and expertise in modern CI/CD and IaC techniques. Compensation includes a competitive salary range of 152,000 to 287,500 USD depending on level and experience, plus equity and benefits.

Qualifications

  • 5+ years of professional experience building and supporting critical services.
  • Experience supporting large-scale HPC clusters with setup, tuning, and troubleshooting.
  • Strong experience crafting large-scale infrastructure platforms.

Responsibilities

  • Own SRE solutions end-to-end from design to operation.
  • Ensure highest uptime and quality of service for internal customers.
  • Conduct capacity management to meet operational needs.

Skills

Infrastructure as Code (IaC)
Python
Kubernetes
HPC clusters
CI/CD techniques
Debugging

Education

B.S. in Computer Science or equivalent

Tools

Slurm
AWS
GCP
OCI

Job description

A leading technology company in Austin, Texas is seeking a Senior SRE to join their Compute Farm team. In this role, you will be responsible for owning SRE solutions and ensuring the highest uptime for missions-critical services. The ideal candidate should have at least 5 years of experience in supporting large-scale HPC clusters and expertise in modern CI/CD and IaC techniques. Compensation includes a competitive salary range of 152,000 to 287,500 USD depending on level and experience, plus equity and benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC SRE: Global Multi-Cloud Reliability
Senior HPC SRE: Global Multi-Cloud Reliability

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Senior HPC SRE: Scale Research Clusters & Uptime
Senior HPC SRE: Scale Research Clusters & Uptime

The Voleon Group • Berkeley (CA)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer: Multi-Cloud & IaC
Senior Site Reliability Engineer: Multi-Cloud & IaC

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 100,000 - 130,000
Competitive salary and benefits
Growth and learning opportunities
Friendly and collaborative team environment
Principal SRE: Hybrid Cloud-Native Reliability Leader
Principal SRE: Hybrid Cloud-Native Reliability Leader

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 146,000 - 221,000
Hybrid Cloud SRE Engineer – Automation & Reliability
Hybrid Cloud SRE Engineer – Automation & Reliability

DryvIQ • United States

Hybrid
USD 120,000 - 160,000
Senior SRE: Build Resilient, Scalable Cloud Systems (Remote)
Senior SRE: Build Resilient, Scalable Cloud Systems (Remote)

Roman Health Pharmacy LLC • New York (NY)

Hybrid
USD 182,000 - 220,000
Full medical, dental, and vision insurance
401(k) with company match
Flexible PTO
+4
Senior HPC SRE Engineer - Automation & Linux Systems
Senior HPC SRE Engineer - Automation & Linux Systems

Omega Enterprise Solutions, LLC • Corridor North (MD)

On-site
USD 100,000 - 130,000
Senior SRE: Build Reliability for Global Bare Metal Cloud
Senior SRE: Build Reliability for Global Bare Metal Cloud

Latitude.sh • United States

Remote
USD 120,000 - 150,000
Senior SRE: Multi-Cloud Reliability for Live Games
Senior SRE: Multi-Cloud Reliability for Live Games

2K • Austin (TX)

On-site
USD 120,000 - 160,000
Senior Cloud SRE – FedRAMP & Multi-Cloud Infra
Senior Cloud SRE – FedRAMP & Multi-Cloud Infra

Palo Alto Networks • Santa Clara (CA)

On-site
USD 120,000 - 200,000