GPU Platform Reliability Engineer

Beam

San Francisco (CA)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Meaningful equity
Health, dental, vision benefits
Fitness stipend
Learning budget
Events in cloud native community

Job summary

Beam is seeking a role-focused engineer to own the health and reliability of our GPU compute fleet in a fast-growing AI inference platform. You will build and own metrics pipelines, alerts, and a unified health view across thousands of GPUs in production.

You will automate deployment debugging, create a scalable firmware telemetry stack, and define the qualification process for new GPUs onboarded to our platform. Join Beam at the ground floor of a rapidly growing startup.

Qualifications

  • Experience with firmware-level diagnostics and telemetry.
  • Proven ability to design health dashboards.
  • Experience with GPU compute hardware.

Responsibilities

  • Own compute fleet health end to end with metrics pipelines and alerting.
  • Automate deployment debugging from detection to triage.
  • Design the GPU qualification platform including burn-in and performance baselining.
  • Own firmware-level telemetry and log collection at scale.

Skills

Hardware fault analysis
Firmware debugging
GPU platforms
Open source

Tools

Telemetry pipelines
Logging at scale
CI/CD

Job description

Beam is seeking a role-focused engineer to own the health and reliability of our GPU compute fleet in a fast-growing AI inference platform. You will build and own metrics pipelines, alerts, and a unified health view across thousands of GPUs in production.

You will automate deployment debugging, create a scalable firmware telemetry stack, and define the qualification process for new GPUs onboarded to our platform. Join Beam at the ground floor of a rapidly growing startup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU SRE & Firmware Telemetry Engineer
GPU SRE & Firmware Telemetry Engineer

Beam • California (MO)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Fitness stipend
+2
Site Reliability Engineer
Site Reliability Engineer

ApplyMint • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 270,000
Competitive salary
Meaningful equity
Health, dental, and vision benefits
+2
Site Reliability Engineer
Site Reliability Engineer

Beam • California (MO)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Fitness stipend
+2
Site Reliability Engineer
Site Reliability Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary
Meaningful equity
Health, dental, vision benefits
+3
Platform Engineer - GPU-Scale Cloud Native
Platform Engineer - GPU-Scale Cloud Native

Beam • California (MO)

On-site
USD 140,000 - 190,000
Competitive salary
Equity
Health benefits
+2
Distributed Systems Engineer — Scale GPU Cloud Platforms
Distributed Systems Engineer — Scale GPU Cloud Platforms

Beam • New York (NY)

On-site
USD 120,000 - 170,000
Competitive salary
Equity
Health, dental, vision
+2
Platform Engineer — Distributed Systems, Kubernetes & GPUs
Platform Engineer — Distributed Systems, Kubernetes & GPUs

Beam • California (MO)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Health benefits
+4
Software Engineer, Platform
Software Engineer, Platform

Beam • California (MO)

On-site
USD 140,000 - 190,000
Competitive salary
Equity
Health benefits
+2
GPU Fleet SRE for AI Inference Platform
GPU Fleet SRE for AI Inference Platform

ApplyMint • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 270,000
Competitive salary
Meaningful equity
Health, dental, and vision benefits
+2
Distributed Systems Engineer
Distributed Systems Engineer

Beam • New York (NY)

On-site
USD 120,000 - 170,000
Competitive salary
Equity
Health, dental, vision
+2