Site Reliability Engineer

ltdglobal

Berkeley (CA)

Hybrid

USD 96,000 - 124,000

Part time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ltdglobal is seeking a sharp, self-motivated Site Reliability Engineer to support a national HPC facility in Berkeley, CA. This 1-year contract assignment offers $80/hr with a possibility of extension based on performance and organizational needs.

You will monitor and triage real-time alerts across compute, storage, network and facility systems; build automation to prevent outages; develop monitoring tools and integrations; and walk the data center floor to maintain power, cooling, and

Qualifications

  • Solid Linux/command-line (SSH) chops.
  • Programming/scripting experience: Python, C, C++, Perl, or Java.
  • Self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems.
  • Comfortable with Owl shift (12am–8am), five days/week, hybrid onsite in Berkeley, CA.
  • Strong cross-team communication and collaboration skills.

Responsibilities

  • Monitor and triage alerts across compute, storage, network, and facility systems in real time.
  • Build automation that prevents issues before outages occur.
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action).
  • Walk the data center floor to keep power, cooling, and environmental systems humming.
  • Coordinate maintenance activities across teams and keep incidents accurately tracked.
  • Dig into complex, ambiguous problems and drive them to resolution.

Skills

Linux CLI
Python
C/C++
Perl
Java
Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
Networking security
Cross-team collaboration

Tools

ServiceNow
ITSM

Job description

Hybrid — Berkeley, CA

1 YearContract Assignment w ith possibility of extension based on performance and organizational needs.

$80/hr

Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.

If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.

What You'll Own
  • Monitor and triage alerts across compute, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution
What You Bring
  • Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience:Python, C, C++, Perl, or Java
  • A self-starter mindset,eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills
Nice to Have
  • Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

On-site
USD 94,000 - 127,000
Site Reliability Engineer Hybrid
Site Reliability Engineer Hybrid

LTD GLOBAL, LLC • Berkeley (CA)

On-site
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
B Site Reliability Engineer Bay Systems Berkeley, California, US $80-80
B Site Reliability Engineer Bay Systems Berkeley, California, US $80-80

Artha Nexgen • Berkeley (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Safety Reliability Engineer II - Supporting large-scale data centers
Safety Reliability Engineer II - Supporting large-scale data centers

BCP Engineers & Consultants • Berkeley (CA)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems Consulting, Inc. (BSC) • Berkeley (CA)

On-site
USD 110,000 - 160,000
SRE for HPC Infra & Automation — Overnight (Contract)
SRE for HPC Infra & Automation — Overnight (Contract)

ltdglobal • Berkeley (CA)

Hybrid
USD 96,000 - 124,000
Night-Shift SRE for High-Impact HPC & Automation
Night-Shift SRE for High-Impact HPC & Automation

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000