Reliability Engineer - Large-Scale AI & HPC Systems

PVH (Tommy Hilfiger/Calvin Klein)

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack is seeking a Data Center Operations professional to own reliability for customer workloads and coordinate across teams. You’ll debug across the full stack and manage incident communications with depth and clarity, driving fixes from production through engineering.

This role suits those with experience supporting HPC/cloud/AI labs and a track record of preventing repeat issues. You will push internal teams to address root causes and maintain SLAs while operating at scale beyond a

Qualifications

  • You’ve supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don’t own.
  • You have written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.

Responsibilities

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer‑facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

Skills

Large-scale compute experience
Distributed systems debugging
Incident communication
Root-cause analysis

Tools

Slurm
Kubernetes
NCCL
InfiniBand / RoCE

Job description

Fluidstack is seeking a Data Center Operations professional to own reliability for customer workloads and coordinate across teams. You’ll debug across the full stack and manage incident communications with depth and clarity, driving fixes from production through engineering.

This role suits those with experience supporting HPC/cloud/AI labs and a track record of preventing repeat issues. You will push internal teams to address root causes and maintain SLAs while operating at scale beyond a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Compute Reliability Engineer — Scale & Incident Leadership
AI Compute Reliability Engineer — Scale & Incident Leadership

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
Senior AI Compute Reliability Engineer
Senior AI Compute Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Senior Data Center Reliability Engineer
Senior Data Center Reliability Engineer

Fluidstack • Austin (TX)

On-site
USD 220,000 - 260,000
Reliability Engineer - AI Infra & Data Centers
Reliability Engineer - AI Infra & Data Centers

Fluidstack • Seattle (WA), New York (NY), San Francisco (CA), Austin (TX)

On-site
USD 120,000 - 180,000
Principal Data Center Reliability Engineer
Principal Data Center Reliability Engineer

Fluidstack • United States

On-site
USD 220,000 - 260,000
Customer Reliability Engineer
Customer Reliability Engineer

fluidstack • San Francisco (CA)

On-site
USD 120,000 - 180,000
Network Reliability Engineer — Automation & AI Tooling
Network Reliability Engineer — Automation & AI Tooling

Fluidstack • Austin (TX)

On-site
USD 208,000 - 269,000
Salary growth potential
Premium health benefits
Data Center Reliability Engineer — Design & Modeling (Equity)
Data Center Reliability Engineer — Design & Modeling (Equity)

FluidStack • United States

On-site
USD 200,000 - 250,000
Stock options
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000