Site Reliability Engineer

Ltd Global

Berkeley (CA)

Hybrid

USD 94,000 - 127,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ltd Global is seeking a sharp, self-motivated SRE to support a national HPC facility in Berkeley, CA. The role covers real-time monitoring, automation, and cross-team coordination to keep compute, storage, and facility systems running smoothly.

The candidate should be comfortable with the Owl shift (12am–8am), have strong Linux/CLI skills, and experience in Python or C-family languages. Hybrid onsite work is required in Berkeley, CA.

Qualifications

  • Comfort working Owl shift (12am-8am), 5 days/week, hybrid onsite in Berkeley, CA.
  • Solid Linux/command-line (SSH) chops.
  • Programming/scripting experience: Python, C, C++, Perl, or Java.
  • Self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems.
  • Network security fundamentals (ACLs, firewalls).
  • Strong cross-team communication and collaboration skills.

Responsibilities

  • Monitor and triage alerts across compute, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs - alerts - action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution

Skills

Linux CLI
Python
C/C++/Java

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
SSH

Job description

Hybrid - Berkeley, CA

1 Year Contract Assignment with possibility of extension based on performance and organizational needs.

$80/hr

Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.

What You'll Own
  • Monitor and triage alerts across compute, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs - alerts - action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution
What You Bring
  • Comfort working Owl shift (12am-8am), 5 days/week, hybrid onsite in Berkeley, CA
  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience:Python, C, C++, Perl, or Java
  • A self-starter mindset,eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills
Nice to Have
  • Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Hybrid
Site Reliability Engineer Hybrid

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Night-Shift SRE for High-Impact HPC & Automation
Night-Shift SRE for High-Impact HPC & Automation

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
SRE/Platform Engineer
SRE/Platform Engineer

Stash Talent Services • Virginia (MN)

Remote
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mikealbert • Cincinnati (OH)

Hybrid
USD 100,000 - 130,000