Compute Fleet Reliability Engineer | GPU/TPU

Jobtailor

San Francisco (CA)

On-site

USD 170,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

FluidStack is seeking a hands-on reliability engineer to own the compute fleet health at scale in production environments, including Kubernetes and bare metal. You will build metrics pipelines, alerting, and a unified health view for GPUs/TPUs, and automate repair workflows from detection to service.

You will design the XPU qualification platform, define good-enough baselines, and own Redfish/BMC tooling and telemetry.

Qualifications

  • Toil as a bug: manual steps in a repair workflow are a backlog item.
  • Strong hardware fault analysis at firmware/silicon level.
  • Walk into the fog, build the map, and explain it to others.
  • Learn quickly in unfamiliar domains; this is valued over prior expertise.
  • Carry a pager: respond to incidents, write postmortems, fix root causes.
  • Fluent with AI tooling and AI coding environments.

Responsibilities

  • Own compute fleet health end to end with metrics pipelines and alerts.
  • Turn repair into automation from detection to service.
  • Design and expand the XPU qualification platform for GPUs/TPUs.
  • Own Redfish and BMC tooling and fleet telemetry.
  • Own reliability, scalability, and operation of the compute fleet at scale.

Skills

Toil as a bug
Hardware intuition
Ambiguity navigation
Rapid learning
Pager/Incident response
AI tooling fluency

Tools

Redfish tooling
BMC tooling
IPMI tooling
Prometheus
Grafana
Go
Python

Job description

FluidStack is seeking a hands-on reliability engineer to own the compute fleet health at scale in production environments, including Kubernetes and bare metal. You will build metrics pipelines, alerting, and a unified health view for GPUs/TPUs, and automate repair workflows from detection to service.

You will design the XPU qualification platform, define good-enough baselines, and own Redfish/BMC tooling and telemetry.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE - Compute & Hyperscale GPU Fleet Reliability
SRE - Compute & Hyperscale GPU Fleet Reliability

Fluidstack • New York (NY)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Equity participation
Retirement plan
+1
Production Engineer, Compute
Production Engineer, Compute

Jobtailor • San Francisco (CA)

On-site
USD 170,000 - 240,000
Production Engineer, Compute - Automate & Scale GPU Fleet
Production Engineer, Compute - Automate & Scale GPU Fleet

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Competitive total compensation
Retirement or pension plan
Health, dental, and vision insurance
+1
Compute Production Lead: GPU Fleet & Automation
Compute Production Lead: GPU Fleet & Automation

Fluidstack • Seattle (WA), New York (NY), Austin (TX), San Francisco (CA)

On-site
USD 180,000 - 280,000
Head of GPU Compute Production Engineering
Head of GPU Compute Production Engineering

Fluidstack • San Francisco (CA)

On-site
USD 269,000 - 335,000
Lead GPU Compute Production Engineering
Lead GPU Compute Production Engineering

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
Reliability Engineer - Large-Scale AI & HPC Systems
Reliability Engineer - Large-Scale AI & HPC Systems

PVH (Tommy Hilfiger/Calvin Klein) • San Francisco (CA)

On-site
USD 120,000 - 180,000
Principal Data Center Reliability Engineer
Principal Data Center Reliability Engineer

Fluidstack • United States

On-site
USD 220,000 - 260,000
Lead Production Engineering for GPU Compute Ops
Lead Production Engineering for GPU Compute Ops

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Compute Systems Engineer — GPU & Server Validation
Compute Systems Engineer — GPU & Server Validation

Fluidstack • Seattle (WA), New York (NY), San Francisco (CA), Austin (TX)

On-site
USD 140,000 - 190,000