Night-Shift SRE for High-Impact HPC & Automation

Ltd Global

Berkeley (CA)

Hybrid

USD 94,000 - 127,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ltd Global is seeking a sharp, self-motivated SRE to support a national HPC facility in Berkeley, CA. The role covers real-time monitoring, automation, and cross-team coordination to keep compute, storage, and facility systems running smoothly.

The candidate should be comfortable with the Owl shift (12am–8am), have strong Linux/CLI skills, and experience in Python or C-family languages. Hybrid onsite work is required in Berkeley, CA.

Qualifications

  • Comfort working Owl shift (12am-8am), 5 days/week, hybrid onsite in Berkeley, CA.
  • Solid Linux/command-line (SSH) chops.
  • Programming/scripting experience: Python, C, C++, Perl, or Java.
  • Self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems.
  • Network security fundamentals (ACLs, firewalls).
  • Strong cross-team communication and collaboration skills.

Responsibilities

  • Monitor and triage alerts across compute, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs - alerts - action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution

Skills

Linux CLI
Python
C/C++/Java

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
SSH

Job description

Ltd Global is seeking a sharp, self-motivated SRE to support a national HPC facility in Berkeley, CA. The role covers real-time monitoring, automation, and cross-team coordination to keep compute, storage, and facility systems running smoothly.

The candidate should be comfortable with the Owl shift (12am–8am), have strong Linux/CLI skills, and experience in Python or C-family languages. Hybrid onsite work is required in Berkeley, CA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000
Night-Shift SRE for HPC & Data Center Ops
Night-Shift SRE for HPC & Data Center Ops

JobCubby • Berkeley (CA), Northern (KY)

Hybrid
USD 96,000 - 124,000
Onsite Night SRE for HPC & Data Center Ops
Onsite Night SRE for HPC & Data Center Ops

Essnova Solutions, Inc. • Berkeley (CA)

On-site
USD 91,000 - 129,000
Site Reliability Engineer - HPC & 24/7 Monitoring
Site Reliability Engineer - HPC & 24/7 Monitoring

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer Hybrid
Site Reliability Engineer Hybrid

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Night-Shift HPC Systems Lead - Remote
Night-Shift HPC Systems Lead - Remote

5C Group • Springfield (OH)

On-site
USD 115,000 - 145,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Remote Night-Shift SRE — Cloud Infra Reliability
Remote Night-Shift SRE — Cloud Infra Reliability

Peraton • Reston (VA)

On-site
USD 104,000 - 166,000
Night-Shift Systems Admin Lead - HPC & GPU Clusters
Night-Shift Systems Admin Lead - HPC & GPU Clusters

5C Group • United States

On-site
USD 115,000 - 145,000