Customer Reliability Engineer

fluidstack

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack is seeking a Data Center Operations professional to own reliability for customer workloads and coordinate across teams. You’ll debug across the full stack and manage incident communications with depth and clarity, driving fixes from production through engineering.

This role suits those with experience supporting HPC/cloud/AI labs and a track record of preventing repeat issues. You will push internal teams to address root causes and maintain SLAs while operating at scale beyond a

Qualifications

  • You’ve supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don’t own.
  • You have written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.

Responsibilities

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer‑facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

Skills

Large-scale compute experience
Distributed systems debugging
Incident communication
Root-cause analysis

Tools

Slurm
Kubernetes
NCCL
InfiniBand / RoCE

Job description

The Data Center Operations Team

Examples of key problems the team is working on:

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 100 GW.
  • Fly the plane while it’s being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don’t inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
Role Scope
  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer‑facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.
What We’re Looking For
  • You’ve supported large‑scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don’t own.
  • You have written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.

Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Seattle (WA), New York (NY), Austin (TX), San Francisco (CA)

On-site
USD 180,000 - 280,000
Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • United States

On-site
USD 220,000 - 260,000
Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • Austin (TX)

On-site
USD 220,000 - 260,000
Compute Engineer, Deployment
Compute Engineer, Deployment

Fluidstack • San Francisco (CA)

On-site
USD 150,000 - 250,000
Compute Deployment Engineer
Compute Deployment Engineer

Fluidstack • New York (NY)

On-site
USD 120,000 - 190,000
Kubernetes-based bare-metal Provisiong
Accelerator platform bring-up
Burn-in and stress harness design
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • San Francisco (CA)

On-site
USD 269,000 - 335,000
Reliability Engineer, Data Center Design
Reliability Engineer, Data Center Design

Fluidstack • Austin (TX)

On-site
USD 200,000 - 250,000
Stock options
Reliability Engineer, Data Center Design
Reliability Engineer, Data Center Design

FluidStack • United States

On-site
USD 200,000 - 250,000
Stock options