Site Reliability Engineer

Beam

California (MO)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health benefits
Fitness stipend
Learning budget
Community events

Job summary

Beam is seeking a hardware-focused engineer to own compute fleet health end to end, building metrics pipelines, alerting, and a unified health view for GPU production. You will automate deployment debugging, design the GPU qualification platform with burn‑in and baselining, and own firmware telemetry and scalable log collection to support repair tooling.

Join a fast-growing pre‑series A company with strong equity, health benefits, and opportunities to contribute across cloud native and

Qualifications

  • Strong hardware intuition, capable of reasoning about firmware and silicon failure modes.
  • Fluent with AI tooling and able to maximize token usage for specs.
  • Experience debugging production issues from triage to post‑mortem.
  • Enthusiasm for developer tools, cloud native technologies, and open source software.

Responsibilities

  • Own compute fleet health end to end, build metrics pipelines, alerting, and a unified health view for GPUs in production.
  • Turn deployment debugging into an automated pipeline, building automation from detection through triage.
  • Design the GPU qualification platform with burn‑in, baselining, and NPI execution for new GPUs.
  • Own firmware telemetry, scale log collection, and the low‑level access layer used by repair tooling.

Skills

Hardware intuition
AI tooling
Production debugging
Cloud native
Open source

Job description

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer‑tool founders, including the founder of Snyk and former CTO of GitHub.

About the Role
  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production.
  • Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage.
  • Design the GPU qualification platform. Burn‑in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production.
  • Own firmware‑level telemetry, log collection at scale, and the low‑level access layer that repair automation and health tooling depend on.
Skills & Experience
  • You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it
  • You're fluent with AI tooling. You aren’t afraid to max‑out your token usage for the right spec.
  • You’re comfortable debugging production issues, from triage to post‑mortem.
  • Enthusiasm for developer tools, cloud native technologies, and open source software
Benefits
  • Competitive salary and meaningful equity
  • Join a fast‑growing pre‑series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary
Meaningful equity
Health, dental, vision benefits
+3
Distributed Systems Engineer
Distributed Systems Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 200,000
Competitive compensation
Equity
Health, dental, vision benefits
+3
Distributed Systems Engineer
Distributed Systems Engineer

Beam • New York (NY)

On-site
USD 120,000 - 170,000
Competitive salary
Equity
Health, dental, vision
+2
Software Engineer, Platform
Software Engineer, Platform

Beam • California (MO)

On-site
USD 140,000 - 190,000
Competitive salary
Equity
Health benefits
+2
Distributed Systems Engineer
Distributed Systems Engineer

Beam • California (MO)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Health benefits
+4
Network Engineer
Network Engineer

Beam • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Health/dental/vision
+2
Network Engineer
Network Engineer

Beam • New York (NY)

On-site
USD 120,000 - 180,000
Health, dental, vision benefits
Equity
Learning budget
GPU Platform Reliability Engineer
GPU Platform Reliability Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary
Meaningful equity
Health, dental, vision benefits
+3
GPU SRE & Firmware Telemetry Engineer
GPU SRE & Firmware Telemetry Engineer

Beam • California (MO)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Fitness stipend
+2
Platform Engineer - GPU-Scale Cloud Native
Platform Engineer - GPU-Scale Cloud Native

Beam • California (MO)

On-site
USD 140,000 - 190,000
Competitive salary
Equity
Health benefits
+2