SRE for HPC Infra & Automation — Overnight (Contract)

ltdglobal

Berkeley (CA)

Hybrid

USD 96,000 - 124,000

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

ltdglobal is seeking a sharp, self-motivated Site Reliability Engineer to support a national HPC facility in Berkeley, CA. This 1-year contract assignment offers $80/hr with a possibility of extension based on performance and organizational needs.

You will monitor and triage real-time alerts across compute, storage, network and facility systems; build automation to prevent outages; develop monitoring tools and integrations; and walk the data center floor to maintain power, cooling, and

Qualifications

  • Solid Linux/command-line (SSH) chops.
  • Programming/scripting experience: Python, C, C++, Perl, or Java.
  • Self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems.
  • Comfortable with Owl shift (12am–8am), five days/week, hybrid onsite in Berkeley, CA.
  • Strong cross-team communication and collaboration skills.

Responsibilities

  • Monitor and triage alerts across compute, storage, network, and facility systems in real time.
  • Build automation that prevents issues before outages occur.
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action).
  • Walk the data center floor to keep power, cooling, and environmental systems humming.
  • Coordinate maintenance activities across teams and keep incidents accurately tracked.
  • Dig into complex, ambiguous problems and drive them to resolution.

Skills

Linux CLI
Python
C/C++
Perl
Java
Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
Networking security
Cross-team collaboration

Tools

ServiceNow
ITSM

Job description

ltdglobal is seeking a sharp, self-motivated Site Reliability Engineer to support a national HPC facility in Berkeley, CA. This 1-year contract assignment offers $80/hr with a possibility of extension based on performance and organizational needs.

You will monitor and triage real-time alerts across compute, storage, network and facility systems; build automation to prevent outages; develop monitoring tools and integrations; and walk the data center floor to maintain power, cooling, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Night-Shift SRE for High-Impact HPC & Automation
Night-Shift SRE for High-Impact HPC & Automation

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000
Site Reliability Engineer
Site Reliability Engineer

ltdglobal • Berkeley (CA)

Hybrid
USD 96,000 - 124,000
Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

On-site
USD 94,000 - 127,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
24/7 Site Reliability Engineer for HPC Data Center
24/7 Site Reliability Engineer for HPC Data Center

BCP Engineers & Consultants • Berkeley (CA)

On-site
USD 110,000 - 170,000
HPC Site Reliability Engineer - 24/7 Onsite Ops
HPC Site Reliability Engineer - 24/7 Onsite Ops

Bay Systems Consulting, Inc. (BSC) • Berkeley (CA)

On-site
USD 110,000 - 160,000
Site Reliability Engineer Hybrid
Site Reliability Engineer Hybrid

LTD GLOBAL, LLC • Berkeley (CA)

On-site
USD 91,000 - 129,000
Software Engineer SRE
Software Engineer SRE

The Mice Groups, Inc. • Austin (TX)

On-site
USD 83,000 - 96,000
SRE — HPC Platform Engineer (Compute & ML)
SRE — HPC Platform Engineer (Compute & ML)

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options and long-term incentives
Bonuses and discretionary incentives
401(k) retirement plan
+5