Senior HPC SRE: Scale Research Clusters & Uptime

The Voleon Group

Berkeley (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Voleon Group in Berkeley, California is looking for a Senior Cluster Site Reliability Engineer (SRE) to enhance their research compute cluster's reliability and performance. This role requires over 5 years of experience in SRE or DevOps roles, familiarity with HPC frameworks, scripting skills, and cloud infrastructure knowledge. You will ensure maximum uptime, resolve urgent issues, and implement systemic improvements. This is a key position in supporting world-class HPC operations for machine learning research.

Qualifications

  • 5+ years of experience in SRE or DevOps roles.
  • Knowledge of HPC and machine learning training systems.
  • Ability to develop scripts in Python, Ruby, etc.
  • Experience with cloud infrastructure (AWS or GCP).

Responsibilities

  • Be a first responder in the event of cluster outages.
  • Ensure a high degree of cluster uptime and track SLAs.
  • Diagnose systemic problems and engineer solutions.
  • Develop robust metrics and observability for cluster health.
  • Assist in forecasting cluster growth.

Skills

SRE or DevOps experience
HPC/batch compute frameworks
Scripting (Python, Ruby)
Infrastructure-as-code tools
Cloud infrastructure (AWS, GCP)
Modern observability stacks
Distributed storage technologies
System engineer mindset

Education

Bachelor degree in computer science

Tools

Terraform
Ansible
Prometheus
Grafana
Kubernetes
Docker
Slurm

Job description

The Voleon Group in Berkeley, California is looking for a Senior Cluster Site Reliability Engineer (SRE) to enhance their research compute cluster's reliability and performance. This role requires over 5 years of experience in SRE or DevOps roles, familiarity with HPC frameworks, scripting skills, and cloud infrastructure knowledge. You will ensure maximum uptime, resolve urgent issues, and implement systemic improvements. This is a key position in supporting world-class HPC operations for machine learning research.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cluster Site Reliability Engineer
Senior Cluster Site Reliability Engineer

The Voleon Group • Berkeley (CA)

On-site
USD 120,000 - 150,000
Senior SRE Engineer: Scale, Automate, Uptime
Senior SRE Engineer: Scale, Automate, Uptime

Google • Sunnyvale (CA)

On-site
USD 166,000 - 244,000
Senior Site Reliability Engineer — Scale, Automation & Uptime
Senior Site Reliability Engineer — Scale, Automation & Uptime

Bolt Graphics, Inc. • Sunnyvale (CA)

On-site
USD 145,000 - 165,000
100% covered medical, dental, and vision premiums
Equity - Stock Options
401(k) match
Senior HPC SRE — Kubernetes, Linux, & Performance
Senior HPC SRE — Kubernetes, Linux, & Performance

Artha Nexgen • Oak Ridge (TN), Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior SRE: Scale Global Blockchain Infra
Senior SRE: Scale Global Blockchain Infra

Alpen Labs Inc. • United States

On-site
USD 120,000 - 160,000
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer: Cloud, Kubernetes Uptime
Senior Site Reliability Engineer: Cloud, Kubernetes Uptime

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
SRE: Network & Traffic for Hyper-Scale Infra
SRE: Network & Traffic for Hyper-Scale Infra

TikTok • San Jose (CA)

On-site
USD 136,000 - 360,000
Senior SRE: Scale, Reliability & Observability Leader
Senior SRE: Scale, Reliability & Observability Leader

Brez Technology Inc. • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Private Medical, Dental and Vision Benefits
Retirement Savings plan with matching contributions
Workspace benefits for your home office
+4
Senior SRE, BCM/DGX Cloud - Scale GPU Clusters
Senior SRE, BCM/DGX Cloud - Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits