Customer Reliability Engineer

Fluidstack

Seattle (WA)

On-site

USD 173,000 - 224,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack is building civilization-scale infrastructure for AI. The Data Center Operations Team focuses on running reliable, scalable compute for leading customers, owning SLAs, and coordinating incident response as sites come online and scale from tens to hundreds of gigawatts.

If you have experience supporting large-scale HPC, cloud, or AI lab workloads, can debug distributed systems across layers, and communicate incident updates clearly to customers, you may fit this role.

Qualifications

  • Experience supporting large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • Ability to debug distributed systems across layers.
  • Experience writing incident updates trusted by customers.
  • Proactive in driving fixes with production teams.
  • Bonus: GPU training workloads; InfiniBand or RoCE; Slurm or Kubernetes; NCCL debugging.

Responsibilities

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

Skills

Large-scale compute
Distributed systems debugging
Incident communication
Customer workload reliability
GPU training workloads

Job description

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don\'t share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We\'re singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

The Data Center Operations Team

Examples of key problems the team is working on

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don\'t inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.

Role Scope

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

What We\'re Looking For

  • The below is a starting point. We always make space for exceptional people, so if you don\'t fit this role exactly, tell us where you would.
  • You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don\'t own.
  • You've written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.
  • Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.
We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Compensation Range: $173K - $224K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Compute Engineer, Deployment
Compute Engineer, Deployment

Fluidstack • Austin (TX)

On-site
USD 164,000 - 206,000
Compute Engineer, Deployment
Compute Engineer, Deployment

Fluidstack • Seattle (WA)

On-site
USD 197,000 - 227,000
Infrastructure Delivery Program Lead
Infrastructure Delivery Program Lead

Fluidstack • Austin (TX)

On-site
USD 253,000 - 295,000
Infrastructure Delivery Program Lead
Infrastructure Delivery Program Lead

Fluidstack • Seattle (WA)

On-site
USD 186,000 - 220,000
Production Engineer, IaaS Team Lead
Production Engineer, IaaS Team Lead

Fluidstack • San Francisco (CA)

On-site
USD 225,000 - 284,000
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Seattle (WA), New York (NY), Austin (TX), San Francisco (CA)

On-site
USD 180,000 - 280,000
Infrastructure Delivery Program Lead
Infrastructure Delivery Program Lead

Fluidstack • San Francisco (CA)

On-site
USD 253,000 - 295,000
Logistics Technician
Logistics Technician

Fluidstack • Town of Texas (WI)

On-site
USD 115,000 - 142,000
Equity
Retirement plan
Health insurance
+1
Compute Engineer, Deployment Team Lead
Compute Engineer, Deployment Team Lead

Fluidstack • Seattle (WA)

On-site
USD 166,000 - 206,000