Site Reliability Engineer

Bay Systems Consulting Inc.

Berkeley (CA)

On-site

USD 83,000 - 166,000

Full time

37 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NERSC in Berkeley seeks a Site Reliability Engineer to join the 24x7 Operations Technology Group. The role focuses on proactive health monitoring of HPC systems, automation, and incident management to keep compute power available for DOE research.

As a member of a 24x7 team, you will develop and maintain monitoring pipelines, respond to alerts, and collaborate across groups to ensure reliable, scalable operations for energy research and data analysis.

Qualifications

  • Experience in or willingness to work in a 24/7 onsite data center team.
  • Experience on Linux shell and SSH access.
  • Experience with programming languages such as C, C++, Perl, Java, or Python.
  • Familiarity with Kubernetes, Prometheus/VictoriaMetrics, Alertmanager.
  • Experience with ITSM and ServiceNow is a plus.

Responsibilities

  • Monitor the NERSC HPC facility on a 5-day onsite schedule (midnight-8 am).
  • Respond to alerts from systems, storage, network, and data center systems; triage or escalate.
  • Develop and maintain monitoring tools and automation for routine conditions.
  • Configure and maintain application/tool settings to ensure reliability as demand grows.
  • Collaborate with other groups to coordinate center-wide maintenance and workflows.
  • Provide incident information in the trouble ticketing system for outages and updates.

Skills

Linux shell
Programming languages (C, C++, Perl, 1

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
ServiceNow

Job description

About this position

The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. NERSC provides critical HPC and data systems and support for NERSC’s 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that NERSC’s computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.

The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. NERSC provides critical HPC and data systems and support for NERSC’s 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that NERSC’s computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.

ESSENTIAL DUTIES & RESPONSIBILITIES

Describe the duties/functions essential to performing the job.

  • Works an onsite 5-day weekly schedule consisting of Owl (midnight-8 am) shifts to monitor the NERSC HPC Facility.
  • Review and respond to alerts from computer systems, storage, network, and other data center/facility-related systems by triaging or calling the appropriate on‑call staff.
  • Create appropriate solutions to improve processes, prevent issue recurrence, and automate responses to all routine service conditions.
  • Identify issues and propose solutions that will improve monitoring capabilities or provide better automation for triage.
  • Possess expertise in ServiceNow and its usage to develop and implement customized service management solutions.
  • Respond to alerts from multiple systems to ensure that data collection continues 24/7, providing real time information for diagnoses.
  • Develop and maintain tools within the monitoring pipeline in collaboration with the Operations Team.
  • Create new software to provide alerts and notifications from HPC system APIs into the monitoring pipeline.
  • Builds and maintains application/tool configurations to ensure software runs reliably as data and user demands grow.
  • Collaborate with other groups at NERSC to ensure that communication and workflows are clearly understood.
  • Work closely with other NERSC groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.
  • Perform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure to ensure peak operational efficiency.
  • Provide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents so that workflows and protocols can be appropriately tracked by others.
  • Work on and resolve problems of diverse scope where data analysis requires the evaluation of identifiable factors.
  • Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
  • Work on and resolve complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors.
POSITION REQUIREMENTS
  • Experience in or willingness to work within a 24/7 onsite team environment to support large-scale data centers or critical installations.
  • Experience on Linux shell and working in a command-line (e.g. SSH) environment.
  • Experience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
  • Motivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.
  • Experience with network security: configuring/maintaining ACLs, knowledge of firewalls.
  • Experience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.
  • Good to Have : Practical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.
  • Experience with ServiceNow implementation is a plus.
  • Familiarity with ITSM best practices and an understanding of how to align service lifecycles with business goals is preferred.
Knowledge, Skills & Abilities
  • Strong hands‑on knowledge of the Linux shell and working in a command-line (e.g. SSH) environment.
  • Strong hands‑on knowledge on various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
  • Knowledge of and ability to work on large data communications networks/ Network Protocols and IT infrastructure supporting highly available systems and applications.
  • Strong communication skills and ability to work effectively across multiple technical teams.
  • Good to Have : Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to automate decision‑making, optimize complex workflows, and enhance proactive system monitoring.
Salary Information

$0 - $80Hourly Wage

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
ASKUSR0145930 Site Reliability Engineer (SRE)
ASKUSR0145930 Site Reliability Engineer (SRE)

JobCubby • Berkeley (CA), Northern (KY)

Hybrid
USD 96,000 - 124,000
ASKUSR0145930 Site Reliability Engineer (SRE)
ASKUSR0145930 Site Reliability Engineer (SRE)

Essnova Solutions, Inc. • Berkeley (CA)

On-site
USD 91,000 - 129,000
Site Reliability Engineer - HPC & 24/7 Monitoring
Site Reliability Engineer - HPC & 24/7 Monitoring

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer Hybrid
Site Reliability Engineer Hybrid

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

On-site
USD 94,000 - 127,000
HPC Site Reliability Engineer — Onsite 24/7
HPC Site Reliability Engineer — Onsite 24/7

Bay Systems Consulting Inc. • Berkeley (CA)

On-site
USD 83,000 - 166,000
Systems Engineer III
Systems Engineer III

NISC • Atlanta (GA)

On-site
USD 120,000 - 180,000