Get more replies from employers
Send a job-specific resume in minutes.
Great Organization! in Berkeley, CA is seeking an experienced Site Reliability Engineer to join the Operations Technology team.
This hybrid role supports a national-scale HPC facility, ensuring uptime, reliability, and security through proactive monitoring, automation, and cross-functional collaboration. The engineer will own monitoring and triage, build automation, and develop tooling across the monitoring pipeline, while supporting data center power, cooling, and environmental controls.
Job Description
** Hybrid — Berkeley, CA**
** Assignment: 10/26/2026 – 10/27/2027**
** $80/hr**
Role Summary
As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a national-scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you'll proactively catch issues before they escape, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You'll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption.
What You Own
What You Bring
Nice to Have
Company Description
Great Organization!
Great Organization!