Reliability Engineer, Data Center Design

FluidStack

United States

On-site

USD 200,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options

Job summary

Fluidstack is seeking a Reliability Engineer to join our Data Center Design Team. You will build and maintain reliability models and run Monte Carlo simulations to quantify availability across large-scale data-center projects.

You will own design-gate reliability studies, benchmark against Uptime Institute tiers, and manage external consultants. Expertise in Windchill/ReliaSoft and MEP systems is essential for success.

Qualifications

  • You've personally built reliability block diagrams and fault tree analyses for mission-critical electrical and mechanical systems.
  • You've run Monte Carlo simulations for system availability using PTC Windchill Prediction, ReliaSoft, or equivalent tools, and you know IEEE 493 (Gold Book) and IEEE 3006.5 well enough to defend your failure-rate and repair-time assumptions.
  • You've managed reliability models, component data, and version control inside PTC Windchill as the system of record, not a spreadsheet on the side.
  • You understand data center MEP systems well enough to model them accurately: MV/LV electrical distribution, standby generation, UPS, chilled water plants, CDUs, and building management/controls systems.
  • You default to quantifying risk instead of describing it. You would rather hand someone a P90 availability number than tell them a system should be reliable.
  • You translate technical reliability findings into figures a lease document, SLA, or investor deck can actually use, without losing what the number means.
  • Bonus: PE license. Direct experience benchmarking designs against Uptime Institute Tier III/IV classifications. Liquid or hybrid cooling reliability modeling. Managing outside reliability consultants or engineering firms on a deliverable basis.

Responsibilities

  • Build and maintain reliability block diagrams and fault tree analyses across the full MEP service chain, from utility intake to rack-level IT load, running Monte Carlo simulations of at least 100,000 iterations to produce P50/P90/P95/P99 availability distributions.
  • Own the reliability study at each 30% and 90% design gate for Fluidstack's own templates, and run independent reliability assessments of EPC-proposed, colocation, and acquired-site designs benchmarked against Uptime Institute Tier III/IV classifications.
  • Manage third-party reliability consultants and review their RBD, FTA, and Monte Carlo work for methodology soundness, turning identified single points of failure into prioritized design recommendations before capital is committed.
  • Maintain the reliability model, component data, and full audit trail inside Windchill PLM, and produce availability statements, sensitivity analyses, and FMEA summaries for lease documents, SLAs, and investor materials.
  • Feed live-site failure and repair data from Operations and Commissioning back into the models, track modeled availability against measured uptime across the portfolio, and translate the components driving unavailability into maintenance, sparing, and capital allocation recommendations such as N+1 versus 2N.

Skills

Reliability modeling
Monte Carlo simulations
RBD/FTA analyses
Windchill knowledge
Electrical & mechanical systems
Risk quantification

Tools

PTC Windchill
ReliaSoft

Job description

How We Operate
  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

The Data Center Design Team

Examples of key problems the team is working on

  • Lead the design, development, and execution of 50GW+ of data centers this decade.

  • Drive the next generation of liquid cooling and electrical distribution systems.

  • Develop and scale first-of-a-kind modular data centers.

  • Influence behind-the-meter designs and planning for multi-GW campuses.

Role Scope
  • Build and maintain reliability block diagrams and fault tree analyses across the full MEP service chain, from utility intake to rack-level IT load, running Monte Carlo simulations of at least 100,000 iterations to produce P50/P90/P95/P99 availability distributions.

  • Own the reliability study at each 30% and 90% design gate for Fluidstack's own templates, and run independent reliability assessments of EPC-proposed, colocation, and acquired-site designs benchmarked against Uptime Institute Tier III/IV classifications.

  • Manage third-party reliability consultants and review their RBD, FTA, and Monte Carlo work for methodology soundness, turning identified single points of failure into prioritized design recommendations before capital is committed.

  • Maintain the reliability model, component data, and full audit trail inside Windchill PLM, and produce availability statements, sensitivity analyses, and FMEA summaries for lease documents, SLAs, and investor materials.

  • Feed live-site failure and repair data from Operations and Commissioning back into the models, track modeled availability against measured uptime across the portfolio, and translate the components driving unavailability into maintenance, sparing, and capital allocation recommendations such as N+1 versus 2N.

What We are Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've personally built reliability block diagrams and fault tree analyses for mission-critical electrical and mechanical systems, not modeled them in the abstract.

  • You've run Monte Carlo simulations for system availability using PTC Windchill Prediction, ReliaSoft, or equivalent tools, and you know IEEE 493 (Gold Book) and IEEE 3006.5 well enough to defend your failure-rate and repair-time assumptions.

  • You've managed reliability models, component data, and version control inside PTC Windchill as the system of record, not a spreadsheet on the side.

  • You understand data center MEP systems well enough to model them accurately: MV/LV electrical distribution, standby generation, UPS, chilled water plants, CDUs, and building management/controls systems.

  • You default to quantifying risk instead of describing it. You would rather hand someone a P90 availability number than tell them a system should be reliable.

  • You translate technical reliability findings into figures a lease document, SLA, or investor deck can actually use, without losing what the number means.

  • Bonus: PE license. Direct experience benchmarking designs against Uptime Institute Tier III/IV classifications. Liquid or hybrid cooling reliability modeling. Managing outside reliability consultants or engineering firms on a deliverable basis.

Compensation: $200,000 - $250,000 per year, depending on experience, skills, qualifications, and location. Offers equity in the form of stock options.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer, Data Center Design
Reliability Engineer, Data Center Design

Fluidstack • Austin (TX)

On-site
USD 200,000 - 250,000
Stock options
Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • United States

On-site
USD 220,000 - 260,000
Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • Austin (TX)

On-site
USD 220,000 - 260,000
Reliability Engineer, R&D
Reliability Engineer, R&D

Fluidstack • Seattle (WA), New York (NY), San Francisco (CA), Austin (TX)

On-site
USD 120,000 - 180,000
Geotechnical Engineer, Data Center Design
Geotechnical Engineer, Data Center Design

Fluidstack • Austin (TX)

On-site
USD 200,000 - 250,000
Stock options
Equity pay transparency
Mechanical Engineer, Deployment Engineering
Mechanical Engineer, Deployment Engineering

jobr.pro • Austin (TX)

On-site
USD 200,000 - 250,000
Stock options
Pay transparency
Lead Mechanical Engineer, Capacity Delivery
Lead Mechanical Engineer, Capacity Delivery

Fluidstack • Austin (TX)

On-site
USD 250,000 - 292,000
Lead Mechanical Engineer, Deployment Engineering
Lead Mechanical Engineer, Deployment Engineering

jobr.pro • Austin (TX)

On-site
USD 300,000 - 340,000
Equity in stock options
Lead Mechanical Engineer, Deployment Engineering
Lead Mechanical Engineer, Deployment Engineering

FluidStack • United States

On-site
USD 300,000 - 340,000
Stock options
Electrical Engineer, Deployment Engineering
Electrical Engineer, Deployment Engineering

jobr.pro • Austin (TX)

On-site
USD 200,000 - 250,000
Stock options