Site Reliability Engineer Hybrid

LTD GLOBAL, LLC

Berkeley (CA)

Hybrid

USD 91,000 - 129,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Great Organization! in Berkeley, CA is seeking an experienced Site Reliability Engineer to join the Operations Technology team.

This hybrid role supports a national-scale HPC facility, ensuring uptime, reliability, and security through proactive monitoring, automation, and cross-functional collaboration. The engineer will own monitoring and triage, build automation, and develop tooling across the monitoring pipeline, while supporting data center power, cooling, and environmental controls.

Qualifications

  • Experience with Linux environments and SSH access.
  • Proficient in scripting or programming languages (Python, C, C++, Java).
  • Familiarity with Kubernetes and monitoring tools like Prometheus.

Responsibilities

  • Monitor and triage alerts across computer, storage, network, and facility systems in real time.
  • Build automation that prevents issues before outages.
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action).
  • Walk the data center floor to keep power, cooling, and environmental systems humming.
  • Coordinate maintenance activities across teams and keep incidents accurately tracked.
  • Dig into complex, ambiguous problems and drive them to resolution.

Skills

Linux/SSH
Python scripting
C/C++
Java

Tools

Kubernetes
Prometheus
VictoriaMetrics
Alertmanager

Job description

Job Description

Job Description

** Hybrid — Berkeley, CA**

** Assignment: 10/26/2026 – 10/27/2027**

** $80/hr**

Role Summary

As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a national-scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you'll proactively catch issues before they escape, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You'll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption.

What You Own

  • Monitor and triage alerts across computer, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution

What You Bring

  • Comfort working Owl shift (12am–8am) , 5 days/week, hybrid onsite in Berkeley, CA
  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience — Python, C, C++, Perl, or Java
  • A self-starter mindset — eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills

Nice to Have

  • Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how


Company Description

Great Organization!

Company Description

Great Organization!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

Hybrid
USD 94,000 - 127,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
Systems Analyst 3 529601671
Systems Analyst 3 529601671

LMG Technology Services LLC • Austin (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer III _Hybrid (w2 only )
Site Reliability Engineer III _Hybrid (w2 only )

Prudent Technologies and Consulting, Inc. • Arlington (TX)

Hybrid
USD 110,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000