Night-Shift SRE for HPC & Data Center Ops

JobCubby

Berkeley, Northern (CA, KY)

Hybrid

USD 96,000 - 124,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the NERSC HPC and data environment for the DOE Office of Science.

The role combines Linux systems administration, infrastructure monitoring, incident response, programming, automation, and data center operations. You will work in a 24/7 operations setting, maintaining reliability and performance across large-scale computing resources, with on-site night shifts and a strong focus on automation and

Qualifications

  • 5+ years of relevant professional experience in Site Reliability Engineering or related field.
  • Strong hands-on Linux administration and SSH proficiency.
  • Experience with one or more programming languages (Python, C, C++, Perl, or Java).
  • Experience supporting large-scale IT infrastructure and HPC environments.
  • Familiarity with network protocols, firewalls/ACLs, and incident triage.
  • Ability to work in a 24/7 operational environment and on-site nights.

Responsibilities

  • Monitor HPC systems, storage, networks, and data center infrastructure.
  • Respond to alerts and perform initial triage with on-call teams.
  • Develop automation to improve monitoring and incident response.
  • Collaborate with teams to resolve bottlenecks and maintain reliability.
  • Use ServiceNow for incident management and workflows.
  • Perform regular data center walkthroughs and document outages.

Skills

Linux
Scripting
Python
C
C++
Perl
Java
SSH
On-call coordination

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
ServiceNow

Job description

Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the NERSC HPC and data environment for the DOE Office of Science.

The role combines Linux systems administration, infrastructure monitoring, incident response, programming, automation, and data center operations. You will work in a 24/7 operations setting, maintaining reliability and performance across large-scale computing resources, with on-site night shifts and a strong focus on automation and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Onsite Night SRE for HPC & Data Center Ops
Onsite Night SRE for HPC & Data Center Ops

Essnova Solutions, Inc. • Berkeley (CA)

On-site
USD 91,000 - 129,000
Site Reliability Engineer - HPC & 24/7 Monitoring
Site Reliability Engineer - HPC & 24/7 Monitoring

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Night-Shift SRE for High-Impact HPC & Automation
Night-Shift SRE for High-Impact HPC & Automation

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Remote Night-Shift SRE — Cloud Infra Reliability
Remote Night-Shift SRE — Cloud Infra Reliability

Peraton • Reston (VA)

On-site
USD 104,000 - 166,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
ASKUSR0145930 Site Reliability Engineer (SRE)
ASKUSR0145930 Site Reliability Engineer (SRE)

JobCubby • Berkeley (CA), Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote Night-Shift SRE: AWS/GovCloud Reliability Engineer
Remote Night-Shift SRE: AWS/GovCloud Reliability Engineer

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
ASKUSR0145930 Site Reliability Engineer (SRE)
ASKUSR0145930 Site Reliability Engineer (SRE)

Essnova Solutions, Inc. • Berkeley (CA)

On-site
USD 91,000 - 129,000