AI Compute Reliability Engineer — Scale & Incident Leadership

Fluidstack

Seattle (WA)

On-site

USD 173,000 - 224,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack is building civilization-scale infrastructure for AI. The Data Center Operations Team focuses on running reliable, scalable compute for leading customers, owning SLAs, and coordinating incident response as sites come online and scale from tens to hundreds of gigawatts.

If you have experience supporting large-scale HPC, cloud, or AI lab workloads, can debug distributed systems across layers, and communicate incident updates clearly to customers, you may fit this role.

Qualifications

  • Experience supporting large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • Ability to debug distributed systems across layers.
  • Experience writing incident updates trusted by customers.
  • Proactive in driving fixes with production teams.
  • Bonus: GPU training workloads; InfiniBand or RoCE; Slurm or Kubernetes; NCCL debugging.

Responsibilities

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

Skills

Large-scale compute
Distributed systems debugging
Incident communication
Customer workload reliability
GPU training workloads

Job description

Fluidstack is building civilization-scale infrastructure for AI. The Data Center Operations Team focuses on running reliable, scalable compute for leading customers, owning SLAs, and coordinating incident response as sites come online and scale from tens to hundreds of gigawatts.

If you have experience supporting large-scale HPC, cloud, or AI lab workloads, can debug distributed systems across layers, and communicate incident updates clearly to customers, you may fit this role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Reliability Engineer
Senior AI Compute Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Cloud Infrastructure Engineer — Hyperscale AI Compute
Cloud Infrastructure Engineer — Hyperscale AI Compute

Fluidstack • New York (NY)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
IaaS Platform Lead — Scale Frontier Compute for AI
IaaS Platform Lead — Scale Frontier Compute for AI

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
IaaS Platform Lead — Self-Service Compute for AI
IaaS Platform Lead — Self-Service Compute for AI

Fluidstack • San Francisco (CA)

On-site
USD 180,000 - 260,000
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Customer Reliability Engineer
Customer Reliability Engineer

fluidstack • San Francisco (CA)

On-site
USD 120,000 - 180,000
End-to-End Infra Delivery Lead, AI Compute
End-to-End Infra Delivery Lead, AI Compute

Fluidstack • San Francisco (CA)

On-site
USD 253,000 - 295,000
Capacity Analytics Engineer — AI Compute & Deployment
Capacity Analytics Engineer — AI Compute & Deployment

Fluidstack • San Francisco (CA)

On-site
USD 143,000 - 173,000
Cloud Infrastructure Engineer — Hyperscale AI Compute
Cloud Infrastructure Engineer — Hyperscale AI Compute

Fluidstack • Seattle (WA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy