Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems

Berkeley (CA)

On-site

USD 120,000 - 150,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

National Energy Research Scientific Computing Center (NERSC) in Berkeley seeks a Site Reliability Engineer to join the 24x7 Operations Technology Group. The role focuses on monitoring the HPC facility, triaging alerts, building automation, and coordinating cross-group activities to maintain high availability and security for thousands of scientific users.

The engineer will develop and maintain monitoring tools, respond to incidents, and work with teams across NERSC to ensure smooth, continuous

Qualifications

  • Experience in or willingness to work within a 24/7 onsite team environment.
  • Experience on Linux shell and a command-line environment.
  • Experience with developing tools using languages such as C, C++, Perl, Java, or Python or scripting.
  • Interest in technologies like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, and power utilization.

Responsibilities

  • Monitor the HPC facility during a 5-day onsite schedule (Owl shifts midnights-8am).
  • Triage alerts from systems, storage, network, and related facilities.
  • Automate responses and improve monitoring to prevent recurrences.

Skills

Linux shell
Programming languages
Communication skills
Team collaboration
Self-starter

Education

Bachelor's degree or higher

Tools

ServiceNow
Kubernetes
Prometheus
Alertmanager
VictoriaMetrics

Job description

National Energy Research Scientific Computing Center (NERSC) in Berkeley seeks a Site Reliability Engineer to join the 24x7 Operations Technology Group. The role focuses on monitoring the HPC facility, triaging alerts, building automation, and coordinating cross-group activities to maintain high availability and security for thousands of scientific users.

The engineer will develop and maintain monitoring tools, respond to incidents, and work with teams across NERSC to ensure smooth, continuous

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - HPC & 24/7 Monitoring
Site Reliability Engineer - HPC & 24/7 Monitoring

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Onsite Night SRE for HPC & Data Center Ops
Onsite Night SRE for HPC & Data Center Ops

Essnova Solutions, Inc. • Berkeley (CA)

On-site
USD 91,000 - 129,000
Night-Shift SRE for HPC & Data Center Ops
Night-Shift SRE for HPC & Data Center Ops

JobCubby • Berkeley (CA), Northern (KY)

Hybrid
USD 96,000 - 124,000
Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000
HPC Scientific Support Engineer | User-Focused & AI-Aware
HPC Scientific Support Engineer | User-Focused & AI-Aware

Berkeley Lab • Berkeley (CA)

Hybrid
USD 156,000 - 219,000
Site Reliability Engineer Hybrid
Site Reliability Engineer Hybrid

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
HPC Network Engineer - Platform, Automation & AI
HPC Network Engineer - Platform, Automation & AI

LBL • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Tuition assistance
Holiday shutdown
Parental leave
+1