Principal Operations Engineer, Reliability

Fluidstack

Austin (TX)

On-site

USD 220,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack, based in Austin, TX, is seeking a Data Center Operations professional to own fleet reliability, lead root-cause analyses, and drive corrective actions across multiple sites. The role involves building a failure data pipeline and shaping maintenance strategy to maximize uptime.

You will work with Weibull, Pareto, and FMEA analyses to guide prioritization and improvements. This position emphasizes ownership, speed, and scalable infrastructure for AI at scale.

Qualifications

  • Experience owning reliability for critical infrastructure and moving the availability metric.
  • Ability to lead root-cause analyses to identify real causes.
  • Fluency with failure data: Weibull, Pareto, FMEA used in practice.

Responsibilities

  • Define and drive fleet reliability targets for the data center fleet.
  • Conduct root-cause analyses on incidents and close actions across sites.
  • Build a failure data pipeline and prioritize engineering work from incident history.
  • Set maintenance strategies (RCM/condition-based) for optimal uptime.

Skills

Reliability engineering
Root-cause analysis
Data analysis
Weibull analysis
FMEA

Tools

CMMS analytics
CRE/CMRP

Job description

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We’re singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them – with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization‑scale infrastructure for AI.

How We Operate
  • Extreme ownership. Full autonomy. Own things end to end, often taking on scope outside your core role without being asked to get things done.
  • Velocity. We drive everything forward as fast as possible.
  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
The Data Center Operations Team
Examples of Key Problems the Team is Working on
  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don’t inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
Role Scope
  • Own fleet reliability engineering: define availability targets, measure them honestly, and close the gap.
  • Run root‑cause analysis on the fleet’s worst incidents and drive corrective actions to completion across every site.
  • Build the failure data pipeline, facilities and hardware both, that turns incident history into engineering priorities.
  • Set the maintenance strategy (reliability‑centered, condition‑based) so the fleet spends effort where the failure data says to.
What We’re Looking For
  • The below is a starting point. We always make space for exceptional people, so if you don’t fit this role exactly, tell us where you would.
  • You’ve owned reliability for critical infrastructure and moved the availability number, not just reported it.
  • You’ve led root‑cause analyses that found the real cause, not the convenient one.
  • You work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use, not terms you know.
  • You get corrective actions closed across teams you don’t manage.
  • Bonus: Data‑center or power‑generation reliability. Liquid cooling systems. CMMS analytics. CRE or CMRP certification.

We are committed to pay equity and transparency.

Equal Employment Opportunity Statement

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability, protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Compensation Range

$220K – $260K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • United States

On-site
USD 220,000 - 260,000
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • San Francisco (CA)

On-site
USD 269,000 - 335,000
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
Technical Program Manager, Data Center Operations
Technical Program Manager, Data Center Operations

Fluidstack • Austin (TX)

On-site
USD 200,000 - 270,000
Health, dental, and vision insurance
Generous PTO policy
Retirement or pension plan
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Facility Operator
Facility Operator

Fluidstack • New York (NY)

On-site
USD 120,000 - 200,000
Equity
Pension plan
Health insurance
+3
Technical Program Manager, Data Center Operations
Technical Program Manager, Data Center Operations

Fluidstack • San Francisco (CA)

On-site
USD 200,000 - 270,000
Competitive total compensation package
Retirement or pension plan
Health, dental, and vision insurance
+1
Technical Program Manager, Data Center Operations
Technical Program Manager, Data Center Operations

Fluidstack • New York (NY)

On-site
USD 200,000 - 270,000
Health, dental, and vision insurance
Competitive total compensation package
Generous PTO policy
Production Engineer, IaaS Team Lead
Production Engineer, IaaS Team Lead

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Facilities Production Technical Lead
Facilities Production Technical Lead

Fluidstack • Austin (TX)

On-site
USD 208,000 - 269,000