Senior AI Compute Reliability Engineer

Fluidstack

New York (NY)

On-site

USD 173,000 - 224,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack is building civilization-scale AI compute infrastructure in the United States. We operate data centers, acquiring power, designing, and running them with teams across hardware and software.

You will own reliability for customer workloads, debug across stacks, and communicate incidents with technical depth to customers. Join a fast-moving team focused on delivering scalable, high-performance compute.

Qualifications

  • Experience with large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • Ability to debug distributed systems across multiple layers.
  • Experience writing incident updates trusted by customers.
  • Familiarity with GPU training workloads and related interconnects (InfiniBand, RoCE).

Responsibilities

  • Own reliability for named customer workloads and SLAs.
  • Debug across the full stack from hardware to fabric to scheduler.
  • Manage customer-facing incident communications with depth and clarity.
  • Translate recurring customer pain into production fixes with engineering teams.

Skills

Distributed systems debugging
HPC/AI labs experience
Incident communications
GPU training workloads
Kubernetes
Slurm
NCCL debugging
InfiniBand
RoCE

Tools

Kubernetes
Slurm
NCCL
InfiniBand
RoCE

Job description

Fluidstack is building civilization-scale AI compute infrastructure in the United States. We operate data centers, acquiring power, designing, and running them with teams across hardware and software.

You will own reliability for customer workloads, debug across stacks, and communicate incidents with technical depth to customers. Join a fast-moving team focused on delivering scalable, high-performance compute.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Compute Reliability Engineer — Scale & Incident Leadership
AI Compute Reliability Engineer — Scale & Incident Leadership

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
Reliability Engineer - Large-Scale AI & HPC Systems
Reliability Engineer - Large-Scale AI & HPC Systems

PVH (Tommy Hilfiger/Calvin Klein) • San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior Software Engineer — AI Infra for World-Scale Compute
Senior Software Engineer — AI Infra for World-Scale Compute

Fluidstack • San Francisco (CA)

On-site
USD 150,000 - 250,000
Equity in stock options
Retirement or pension plan
Health, dental, and vision insurance
+1
Cloud Infrastructure Engineer — Hyperscale AI Compute
Cloud Infrastructure Engineer — Hyperscale AI Compute

Fluidstack • New York (NY)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
AI Infrastructure Engineer — Civilization-Scale Compute
AI Infrastructure Engineer — Civilization-Scale Compute

Fluidstack • San Francisco (CA)

On-site
USD 150,000 - 250,000
Health, dental, and vision insurance
Generous PTO policy
Retirement or pension plan
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Senior Data Center Reliability Engineer
Senior Data Center Reliability Engineer

Fluidstack • Austin (TX)

On-site
USD 220,000 - 260,000
Reliability Engineer, R&D
Reliability Engineer, R&D

Fluidstack • Seattle (WA), New York (NY), San Francisco (CA), Austin (TX)

On-site
USD 120,000 - 180,000
Customer Reliability Engineer
Customer Reliability Engineer

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
IaaS Platform Lead — Self-Service Compute for AI
IaaS Platform Lead — Self-Service Compute for AI

Fluidstack • San Francisco (CA)

On-site
USD 180,000 - 260,000