Senior Data Center Reliability Engineer

Fluidstack

Austin (TX)

On-site

USD 220,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack, based in Austin, TX, is seeking a Data Center Operations professional to own fleet reliability, lead root-cause analyses, and drive corrective actions across multiple sites. The role involves building a failure data pipeline and shaping maintenance strategy to maximize uptime.

You will work with Weibull, Pareto, and FMEA analyses to guide prioritization and improvements. This position emphasizes ownership, speed, and scalable infrastructure for AI at scale.

Qualifications

  • Experience owning reliability for critical infrastructure and moving the availability metric.
  • Ability to lead root-cause analyses to identify real causes.
  • Fluency with failure data: Weibull, Pareto, FMEA used in practice.

Responsibilities

  • Define and drive fleet reliability targets for the data center fleet.
  • Conduct root-cause analyses on incidents and close actions across sites.
  • Build a failure data pipeline and prioritize engineering work from incident history.
  • Set maintenance strategies (RCM/condition-based) for optimal uptime.

Skills

Reliability engineering
Root-cause analysis
Data analysis
Weibull analysis
FMEA

Tools

CMMS analytics
CRE/CMRP

Job description

Fluidstack, based in Austin, TX, is seeking a Data Center Operations professional to own fleet reliability, lead root-cause analyses, and drive corrective actions across multiple sites. The role involves building a failure data pipeline and shaping maintenance strategy to maximize uptime.

You will work with Weibull, Pareto, and FMEA analyses to guide prioritization and improvements. This position emphasizes ownership, speed, and scalable infrastructure for AI at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Data Center Reliability Engineer
Principal Data Center Reliability Engineer

Fluidstack • United States

On-site
USD 220,000 - 260,000
Reliability Engineer - Large-Scale AI & HPC Systems
Reliability Engineer - Large-Scale AI & HPC Systems

PVH (Tommy Hilfiger/Calvin Klein) • San Francisco (CA)

On-site
USD 120,000 - 180,000
Reliability Engineer - AI Infra & Data Centers
Reliability Engineer - AI Infra & Data Centers

Fluidstack • Seattle (WA), New York (NY), San Francisco (CA), Austin (TX)

On-site
USD 120,000 - 180,000
Network Reliability Engineer — Automation & AI Tooling
Network Reliability Engineer — Automation & AI Tooling

Fluidstack • Austin (TX)

On-site
USD 208,000 - 269,000
Salary growth potential
Premium health benefits
Senior AI Compute Reliability Engineer
Senior AI Compute Reliability Engineer

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Data Center Operations Lead
Data Center Operations Lead

Fluidstack • Lubbock (TX)

On-site
USD 120,000 - 180,000
Health, dental, and vision insurance
Retirement plan
Generous PTO policy
+1
Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • United States

On-site
USD 220,000 - 260,000
Data Center Reliability Engineer — Design & Modeling (Equity)
Data Center Reliability Engineer — Design & Modeling (Equity)

FluidStack • United States

On-site
USD 200,000 - 250,000
Stock options
Facilities Production Engineering Lead
Facilities Production Engineering Lead

Fluidstack • Austin (TX)

On-site
USD 208,000 - 269,000
Hyperscale Data Center Operations Manager
Hyperscale Data Center Operations Manager

Apply4itjobs • Houston (TX)

On-site
USD 90,000 - 120,000
Health, dental, and vision insurance
Generous PTO policy
Retirement plan / 401(k)