SRE - Compute & Hyperscale GPU Fleet Reliability

Fluidstack

New York (NY)

On-site

USD 175,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision insurance
Equity participation
Retirement plan
Generous PTO

Job summary

Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling.

The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation.

Qualifications

  • Toil treated as a bug; repairs workflow are not for one-off efforts.
  • Experience reasoning about hardware failure modes at firmware or silicon level.
  • Ability to navigate ambiguity and map complex systems quickly.
  • Rapid learning in unfamiliar domains with strong competence.
  • Ability to handle incidents, write postmortems, and fix systemic causes.
  • Familiar with AI tooling, LLM APIs, MCP servers, agentic frameworks.
  • Experience shipping production automation used by other teams.
  • Bonus: familiarity with BMC/Redfish/IPMI, GPU qualification, and orchestration tools.

Responsibilities

  • Own compute fleet health end-to-end with observability and fleet-scale tooling.
  • Turn deployment and repair into scalable pipelines; automate triage, parts management, and return-to-service.
  • Design and expand the GPU qualification platform for new generations.
  • Own Redfish and BMC tooling for firmware telemetry and fleet-access layers.
  • Ensure reliability, scalability, and operation of the GPU compute fleet at scale.

Skills

Toil-as-bug
Hardware intuition
Ambiguity tolerance
Fast learner
Pager incident handling
AI tooling fluency
Production automation
RMA automation (bonus)

Tools

Go
Python

Job description

Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling.

The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of GPU Compute Production Engineering
Head of GPU Compute Production Engineering

Fluidstack • San Francisco (CA)

On-site
USD 269,000 - 335,000
Lead GPU Compute Production Engineering
Lead GPU Compute Production Engineering

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
Compute Fleet Reliability Engineer | GPU/TPU
Compute Fleet Reliability Engineer | GPU/TPU

Jobtailor • San Francisco (CA)

On-site
USD 170,000 - 240,000
Compute Production Lead: GPU Fleet & Automation
Compute Production Lead: GPU Fleet & Automation

Fluidstack • Seattle (WA), New York (NY), Austin (TX), San Francisco (CA)

On-site
USD 180,000 - 280,000
Production Engineer, Compute - Automate & Scale GPU Fleet
Production Engineer, Compute - Automate & Scale GPU Fleet

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Competitive total compensation
Retirement or pension plan
Health, dental, and vision insurance
+1
Lead Production Engineering for GPU Compute Ops
Lead Production Engineering for GPU Compute Ops

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
SRE - 24/7 GPU/Kubernetes Reliability & Equity
SRE - 24/7 GPU/Kubernetes Reliability & Equity

NVIDIA Gruppe • Town of Texas (WI)

On-site
USD 168,000 - 334,000
Senior SRE: 24/7 GPU & Kubernetes Reliability
Senior SRE: 24/7 GPU & Kubernetes Reliability

Nvidia Corporation in • Austin (TX)

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior SRE, AI Infrastructure & GPU Fleet Reliability
Senior SRE, AI Infrastructure & GPU Fleet Reliability

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
GPU Infrastructure Engineer — Scale & Automate Massive Compute
GPU Infrastructure Engineer — Scale & Automate Massive Compute

Fluidstack • Seattle (WA)

On-site
USD 175,000 - 300,000
Competitive total compensation package
Retirement or pension plan
Health, dental, and vision insurance
+1