Data Center Hardware Quality & Reliability Engineer

OpenAI

United States

Remote

USD 150,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OpenAI seeks a reliability-focused engineer to own the end-to-end hardware quality loop for data-center infrastructure. You will turn field failures into quantified risk, drive fast containment, and ensure verified root cause analysis to improve manufacturing-test coverage and spare forecasting.

The role requires blending practical hardware/system understanding with reliability engineering and data fluency, while leading cross-functional technical teams at data-center scale.

Qualifications

  • Hardware and systems understanding relevant to data-center scale.
  • Experience driving reliability through metrics and data.
  • Ability to lead cross-functional teams in technical settings.

Responsibilities

  • Build and govern the field-quality data model across telemetry and related data.
  • Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF and forecast accuracy with denominators and uncertainty.
  • Provide fleet-level and micro views to detect shifts and bound populations.
  • Lead systemic field-failure triage, containment, failure analysis, and CAPA processes.
  • Develop life-data, reliability-growth, and spare-demand projections by product and geography.
  • Collaborate with MQE and NPI to translate field mechanisms into manufacturing-test coverage and controls.
  • Verify upstream changes reduce field recurrence and improve field performance.
  • Define FA standards, data contracts, scorecards, escalation paths, and closure evidence.
  • Provide inputs on serviceability, TCO, and spares without owning inventory or procurement.
  • Create executive decision packages with risk, options, and recommendations.

Skills

Hardware understanding
Reliability engineering
Data fluency
Cross-functional leadership

Education

Bachelor's degree in Engineering or related field

Tools

SQL
Python
Excel

Job description

About The Role Own the end-to-end data-center hardware quality and reliability loop for OpenAI's 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence. The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.

Key Responsibilities
  • Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.
  • Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.
  • Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.
  • Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.
  • Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age.
  • Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates.
  • Close the loop by verifying whether upstream changes reduce field recurrence.
  • Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence.
  • Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement.
  • Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation.
  • Run the cross-functional reliability council and, as the team grows, mentor
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Center Hardware Reliability Engineer
Senior Data Center Hardware Reliability Engineer

Gohyred • San Francisco (CA)

On-site
USD 180,000 - 280,000
Lead Data Center Hardware Reliability Engineer
Lead Data Center Hardware Reliability Engineer

OpenAI • United States

Remote
USD 150,000 - 210,000
Senior Data Center Hardware Reliability Engineer
Senior Data Center Hardware Reliability Engineer

OpenAI, Inc. • San Francisco (CA)

On-site
USD 226,000 - 285,000
Relocation support
Daily meals in our offices
Learning and development stipend
+1
Data Center Hardware Quality & Reliability Engineer
Data Center Hardware Quality & Reliability Engineer

Gohyred • San Francisco (CA)

On-site
USD 180,000 - 280,000
Reliability Engineer
Reliability Engineer

NextGenEnergyJobs • Austin (TX), Northern (KY)

Hybrid
USD 110,000 - 170,000
Advanced Packaging Reliability Engineer
Advanced Packaging Reliability Engineer

OpenAI • United States

Remote
USD 130,000 - 170,000
Hardware Operations & Fleet Reliability Engineer
Hardware Operations & Fleet Reliability Engineer

OpenAI • Seattle (WA)

On-site
USD 150,000 - 200,000
Reliability Engineer
Reliability Engineer

Speedrun Talent Network • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 190,000
Competitive salary & equity options
Comprehensive benefits package
relocation assistance
+1
Reliability Engineer
Reliability Engineer

Exowatt • Austin (TX)

On-site
USD 110,000 - 165,000
Relocation assistance
Competitive salary and equity options
Comprehensive benefits package
Senior Hardware Reliability Engineer — Data Center Systems
Senior Hardware Reliability Engineer — Data Center Systems

Meta • Fremont (CA)

On-site
USD 144,000 - 204,000