SRE - Compute & Hyperscale GPU Fleet Reliability

Fluidstack

New York (NY)

On-site

USD 175,000 - 300,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision insurance
Equity participation
Retirement plan
Generous PTO

Job summary

Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling.

The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation.

Qualifications

  • Toil treated as a bug; repairs workflow are not for one-off efforts.
  • Experience reasoning about hardware failure modes at firmware or silicon level.
  • Ability to navigate ambiguity and map complex systems quickly.
  • Rapid learning in unfamiliar domains with strong competence.
  • Ability to handle incidents, write postmortems, and fix systemic causes.
  • Familiar with AI tooling, LLM APIs, MCP servers, agentic frameworks.
  • Experience shipping production automation used by other teams.
  • Bonus: familiarity with BMC/Redfish/IPMI, GPU qualification, and orchestration tools.

Responsibilities

  • Own compute fleet health end-to-end with observability and fleet-scale tooling.
  • Turn deployment and repair into scalable pipelines; automate triage, parts management, and return-to-service.
  • Design and expand the GPU qualification platform for new generations.
  • Own Redfish and BMC tooling for firmware telemetry and fleet-access layers.
  • Ensure reliability, scalability, and operation of the GPU compute fleet at scale.

Skills

Toil-as-bug
Hardware intuition
Ambiguity tolerance
Fast learner
Pager incident handling
AI tooling fluency
Production automation
RMA automation (bonus)

Tools

Go
Python

Job description

Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling.

The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE, GPU Compute Fleet—Automation & Reliability
SRE, GPU Compute Fleet—Automation & Reliability

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Generous PTO policy
Retirement or pension plan
Head of GPU Compute Production Engineering
Head of GPU Compute Production Engineering

Fluidstack • San Francisco (CA)

On-site
USD 269,000 - 335,000
Lead GPU Compute Production Engineering
Lead GPU Compute Production Engineering

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
GPU Compute Engineer — Fleet Reliability & Automation
GPU Compute Engineer — Fleet Reliability & Automation

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
Compute Production Lead: GPU Fleet & Automation
Compute Production Lead: GPU Fleet & Automation

Fluidstack • Seattle (WA), New York (NY), Austin (TX), San Francisco (CA)

On-site
USD 180,000 - 280,000
Lead Production Engineering for GPU Compute Ops
Lead Production Engineering for GPU Compute Ops

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Production Engineer, Compute - Automate & Scale GPU Fleet
Production Engineer, Compute - Automate & Scale GPU Fleet

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Competitive total compensation
Retirement or pension plan
Health, dental, and vision insurance
+1
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior SRE: GPU Fleet Orchestration & Auto-Scaling
Senior SRE: GPU Fleet Orchestration & Auto-Scaling

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE — GPU Fleet Orchestration & Autoscaling
Senior SRE — GPU Fleet Orchestration & Autoscaling

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000