The Voleon Group in Berkeley, California is looking for a Senior Cluster Site Reliability Engineer (SRE) to enhance their research compute cluster's reliability and performance. This role requires over 5 years of experience in SRE or DevOps roles, familiarity with HPC frameworks, scripting skills, and cloud infrastructure knowledge. You will ensure maximum uptime, resolve urgent issues, and implement systemic improvements. This is a key position in supporting world-class HPC operations for machine learning research.
Qualifications
5+ years of experience in SRE or DevOps roles.
Knowledge of HPC and machine learning training systems.
Ability to develop scripts in Python, Ruby, etc.
Experience with cloud infrastructure (AWS or GCP).
Responsibilities
Be a first responder in the event of cluster outages.
Ensure a high degree of cluster uptime and track SLAs.
Diagnose systemic problems and engineer solutions.
Develop robust metrics and observability for cluster health.
Assist in forecasting cluster growth.
Skills
SRE or DevOps experience
HPC/batch compute frameworks
Scripting (Python, Ruby)
Infrastructure-as-code tools
Cloud infrastructure (AWS, GCP)
Modern observability stacks
Distributed storage technologies
System engineer mindset
Education
Bachelor degree in computer science
Tools
Terraform
Ansible
Prometheus
Grafana
Kubernetes
Docker
Slurm
Job description
The Voleon Group in Berkeley, California is looking for a Senior Cluster Site Reliability Engineer (SRE) to enhance their research compute cluster's reliability and performance. This role requires over 5 years of experience in SRE or DevOps roles, familiarity with HPC frameworks, scripting skills, and cloud infrastructure knowledge. You will ensure maximum uptime, resolve urgent issues, and implement systemic improvements. This is a key position in supporting world-class HPC operations for machine learning research.