HPC Site Reliability Engineer — Onsite 24/7

Bay Systems Consulting Inc.

Berkeley (CA)

On-site

USD 83,000 - 166,000

Full time

32 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NERSC in Berkeley seeks a Site Reliability Engineer to join the 24x7 Operations Technology Group. The role focuses on proactive health monitoring of HPC systems, automation, and incident management to keep compute power available for DOE research.

As a member of a 24x7 team, you will develop and maintain monitoring pipelines, respond to alerts, and collaborate across groups to ensure reliable, scalable operations for energy research and data analysis.

Qualifications

  • Experience in or willingness to work in a 24/7 onsite data center team.
  • Experience on Linux shell and SSH access.
  • Experience with programming languages such as C, C++, Perl, Java, or Python.
  • Familiarity with Kubernetes, Prometheus/VictoriaMetrics, Alertmanager.
  • Experience with ITSM and ServiceNow is a plus.

Responsibilities

  • Monitor the NERSC HPC facility on a 5-day onsite schedule (midnight-8 am).
  • Respond to alerts from systems, storage, network, and data center systems; triage or escalate.
  • Develop and maintain monitoring tools and automation for routine conditions.
  • Configure and maintain application/tool settings to ensure reliability as demand grows.
  • Collaborate with other groups to coordinate center-wide maintenance and workflows.
  • Provide incident information in the trouble ticketing system for outages and updates.

Skills

Linux shell
Programming languages (C, C++, Perl, 1

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
ServiceNow

Job description

NERSC in Berkeley seeks a Site Reliability Engineer to join the 24x7 Operations Technology Group. The role focuses on proactive health monitoring of HPC systems, automation, and incident management to keep compute power available for DOE research.

As a member of a 24x7 team, you will develop and maintain monitoring pipelines, respond to alerts, and collaborate across groups to ensure reliable, scalable operations for energy research and data analysis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - HPC & 24/7 Monitoring
Site Reliability Engineer - HPC & 24/7 Monitoring

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Onsite Night SRE for HPC & Data Center Ops
Onsite Night SRE for HPC & Data Center Ops

Essnova Solutions, Inc. • Berkeley (CA)

On-site
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Night-Shift SRE for HPC & Data Center Ops
Night-Shift SRE for HPC & Data Center Ops

JobCubby • Berkeley (CA), Northern (KY)

Hybrid
USD 96,000 - 124,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems Consulting Inc. • Berkeley (CA)

On-site
USD 83,000 - 166,000
Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

On-site
USD 94,000 - 127,000
Night-Shift SRE for High-Impact HPC & Automation
Night-Shift SRE for High-Impact HPC & Automation

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000