Senior Data Center Hardware Reliability Engineer

Gohyred

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OpenAI seeks a senior hardware quality and reliability leader to own the end-to-end data-center hardware quality loop for OpenAI’s 3P infrastructure and 1P platforms. You will turn field failures into quantified risk, drive containment, root-cause verification, and improve manufacturing-test coverage.

The role requires practical hardware/system understanding, reliability engineering expertise, data fluency, and cross-functional leadership at data-center scale, with collaboration across MQE and

Qualifications

  • 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure.
  • 3+ years owning field-failure, RMA, or CAPA outcomes.
  • Working proficiency with SQL and Python/R or equivalent analytics tools.

Responsibilities

  • Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s infrastructure and platforms.
  • Define key metrics (AFR, MTBF/MTTR, DPPM) with clear denominators and uncertainty.
  • Lead systemic field-failure triage, containment, and corrective-action verification across fleets.

Skills

SQL
Python/R
Analytics tools
Cross-functional leadership

Education

BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience
MS preferred

Tools

Linux
IPMI
Redfish
BMC

Job description

OpenAI seeks a senior hardware quality and reliability leader to own the end-to-end data-center hardware quality loop for OpenAI’s 3P infrastructure and 1P platforms. You will turn field failures into quantified risk, drive containment, root-cause verification, and improve manufacturing-test coverage.

The role requires practical hardware/system understanding, reliability engineering expertise, data fluency, and cross-functional leadership at data-center scale, with collaboration across MQE and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Center Hardware Reliability Engineer
Senior Data Center Hardware Reliability Engineer

OpenAI, Inc. • San Francisco (CA)

On-site
USD 226,000 - 285,000
Relocation support
Daily meals in our offices
Learning and development stipend
+1
Lead Data Center Hardware Reliability Engineer
Lead Data Center Hardware Reliability Engineer

OpenAI • United States

Remote
USD 150,000 - 210,000
Data Center Hardware Quality & Reliability Engineer
Data Center Hardware Quality & Reliability Engineer

OpenAI • United States

Remote
USD 150,000 - 210,000
Senior Datacenter Hardware Operations Lead
Senior Datacenter Hardware Operations Lead

OpenAI • United States

On-site
USD 86,400 - 228,000
Lead Datacenter Hardware Technician
Lead Datacenter Hardware Technician

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Data Center Hardware Quality & Reliability Engineer
Data Center Hardware Quality & Reliability Engineer

Gohyred • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Hardware Reliability Engineer — Data Center Systems
Senior Hardware Reliability Engineer — Data Center Systems

Meta • Fremont (CA)

On-site
USD 144,000 - 204,000
Hardware Operations & Fleet Reliability Engineer
Hardware Operations & Fleet Reliability Engineer

OpenAI • Seattle (WA)

On-site
USD 150,000 - 200,000
Datacenter Manufacturing Quality Engineer - Global Infra
Datacenter Manufacturing Quality Engineer - Global Infra

OpenAI • United States

Remote
USD 120,000 - 190,000
Lead AI Packaging Reliability Engineer
Lead AI Packaging Reliability Engineer

OpenAI • United States

Remote
USD 130,000 - 170,000