Site Reliability Engineer

Beam

San Francisco (CA)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Meaningful equity
Health, dental, vision benefits
Fitness stipend
Learning budget
Events in cloud native community

Job summary

Beam is seeking a role-focused engineer to own the health and reliability of our GPU compute fleet in a fast-growing AI inference platform. You will build and own metrics pipelines, alerts, and a unified health view across thousands of GPUs in production.

You will automate deployment debugging, create a scalable firmware telemetry stack, and define the qualification process for new GPUs onboarded to our platform. Join Beam at the ground floor of a rapidly growing startup.

Qualifications

  • Experience with firmware-level diagnostics and telemetry.
  • Proven ability to design health dashboards.
  • Experience with GPU compute hardware.

Responsibilities

  • Own compute fleet health end to end with metrics pipelines and alerting.
  • Automate deployment debugging from detection to triage.
  • Design the GPU qualification platform including burn-in and performance baselining.
  • Own firmware-level telemetry and log collection at scale.

Skills

Hardware fault analysis
Firmware debugging
GPU platforms
Open source

Tools

Telemetry pipelines
Logging at scale
CI/CD

Job description

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

About The Role
  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production.
  • Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage.
  • Design the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production.
  • Own firmware-level telemetry, log collection at scale, and the low-level access layer that repair automation and health tooling depend on.
Skills & Experience
  • You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
  • You're fluent with AI tooling. You aren’t afraid to max-out your token usage for the right spec.
  • You’re comfortable debugging production issues, from triage to post-mortem.
  • Enthusiasm for developer tools, cloud native technologies, and open source software.
Benefits
  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ApplyMint • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 270,000
Competitive salary
Meaningful equity
Health, dental, and vision benefits
+2
Site Reliability Engineer
Site Reliability Engineer

Beam • California (MO)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Fitness stipend
+2
Software Engineer, Platform
Software Engineer, Platform

Beam • California (MO)

On-site
USD 140,000 - 190,000
Competitive salary
Equity
Health benefits
+2
Distributed Systems Engineer
Distributed Systems Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 200,000
Competitive compensation
Equity
Health, dental, vision benefits
+3
Distributed Systems Engineer
Distributed Systems Engineer

Beam • New York (NY)

On-site
USD 120,000 - 170,000
Competitive salary
Equity
Health, dental, vision
+2
Distributed Systems Engineer
Distributed Systems Engineer

Beam • California (MO)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Health benefits
+4
Network Engineer
Network Engineer

Beam • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Health/dental/vision
+2
Network Engineer
Network Engineer

Beam • New York (NY)

On-site
USD 120,000 - 180,000
Health, dental, vision benefits
Equity
Learning budget
GPU Platform Reliability Engineer
GPU Platform Reliability Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary
Meaningful equity
Health, dental, vision benefits
+3
GPU SRE & Firmware Telemetry Engineer
GPU SRE & Firmware Telemetry Engineer

Beam • California (MO)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Fitness stipend
+2