24/7 Site Reliability Engineer for HPC Data Center

BCP Engineers & Consultants

Berkeley (CA)

On-site

USD 110,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

BCP Engineers & Consultants is seeking a Site Reliability Engineer to support a scientific computing center on a 24x7 on-site Operations Technology Group. You will monitor, maintain, and secure the computing environment to maximize uptime for researchers and workloads.

The role requires hands-on Linux, scripting, and automation skills, plus collaboration across IT, facilities, and HPC teams. Experience with Kubernetes, Prometheus, ServiceNow, and data-center tooling is highly valued.

Qualifications

  • Experience in or willingness to work within a 24/7 onsite team environment supporting large-scale data centers or critical installations.
  • Strong hands-on experience with the Linux shell and working in a command-line environment (SSH).
  • Experience developing tools using programming languages such as C, C++, Perl, Java, or Python, or a scripting language, with knowledge of standard software development practices.
  • Motivated self-starter able to learn technologies that improve data center management, such as Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.
  • Experience with network security, including configuring/maintaining ACLs and knowledge of firewalls.
  • Knowledge of large data communications networks, network protocols, and IT infrastructure supporting highly available systems and applications.
  • Experience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.
  • Strong communication skills and ability to work effectively across multiple technical teams.

Responsibilities

  • Review and respond to alerts from computer systems, storage, network, and other data center/facility-related systems, triaging or engaging appropriate on-call staff.
  • Respond to alerts from multiple systems to ensure data collection continues 24/7, providing real-time information for diagnoses.
  • Perform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure.
  • Provide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents.

Skills

Linux shell
C/C++
Python
Java
Perl
Scripting
Kubernetes
Monitoring tools
Networking basics
IT security basics

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
ServiceNow

Job description

BCP Engineers & Consultants is seeking a Site Reliability Engineer to support a scientific computing center on a 24x7 on-site Operations Technology Group. You will monitor, maintain, and secure the computing environment to maximize uptime for researchers and workloads.

The role requires hands-on Linux, scripting, and automation skills, plus collaboration across IT, facilities, and HPC teams. Experience with Kubernetes, Prometheus, ServiceNow, and data-center tooling is highly valued.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Site Reliability Engineer - 24/7 Onsite Ops
HPC Site Reliability Engineer - 24/7 Onsite Ops

Bay Systems Consulting, Inc. (BSC) • Berkeley (CA)

On-site
USD 110,000 - 160,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer - HPC & 24/7 Monitoring
Site Reliability Engineer - HPC & 24/7 Monitoring

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
24/7 HPC SRE - Automation & Monitoring
24/7 HPC SRE - Automation & Monitoring

Artha Nexgen • Berkeley (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems Consulting, Inc. (BSC) • Berkeley (CA)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

ltdglobal • Berkeley (CA)

Hybrid
USD 96,000 - 124,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
B Site Reliability Engineer Bay Systems Berkeley, California, US $80-80
B Site Reliability Engineer Bay Systems Berkeley, California, US $80-80

Artha Nexgen • Berkeley (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Site Reliability Engineer for HPC & Control Systems
Senior Site Reliability Engineer for HPC & Control Systems

CT19 • Massachusetts

On-site
USD 140,000 - 210,000