Hardware Operations & Fleet Reliability Engineer

OpenAI

Seattle (WA)

On-site

USD 150,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

OpenAI seeks a Hardware Operations Engineer to serve as the technical authority for hardware reliability and fleet health at a flagship AI campus. You will operate at the intersection of hardware operations, sustaining engineering, and fleet reliability, partnering with CSPs, fleet-health engineers, hardware teams, and OEMs to diagnose and resolve production hardware issues.

The role drives RCA investigations, reliability improvement initiatives, lifecycle programs, and operational readiness,

Qualifications

  • 8+ years of experience supporting large-scale datacenter hardware infrastructure, with senior technician, sustaining engineering, or leadership exposure.
  • Deep expertise with server platforms, GPU systems, storage, rack integration, and datacenter hardware architecture.
  • Strong experience diagnosing complex hardware failures and leading repairs in production environments.
  • Experience conducting root cause analysis and driving long-term corrective actions.
  • Strong understanding of hardware reliability engineering principles and fleet-health management.
  • Proven ability to partner across engineering, operations, manufacturing, and vendor organizations.
  • Comfortable operating independently in high-priority production environments with significant operational responsibility.
  • Excellent written and verbal communication skills with the ability to influence technical and operational decisions.
  • Experience developing operational processes, maintenance standards, and technical documentation.

Responsibilities

  • Drive technical triage and resolution of complex hardware failures impacting production systems.
  • Partner with Fleet Health Engineering to investigate recurring hardware issues, identify failure patterns, and improve fleet reliability.
  • Lead root cause analysis (RCA) efforts for critical hardware incidents and develop corrective and preventive action plans.
  • Collaborate with Cloud Service Provider operations teams and OEM vendors to coordinate repairs, replacements, upgrades, and hardware lifecycle activities.
  • Establish and continuously improve hardware maintenance procedures, operational runbooks, and troubleshooting standards.
  • Analyze hardware failure trends and operational metrics to identify reliability risks and improvement opportunities.
  • Support new hardware introductions, validation activities, and production readiness reviews.
  • Coordinate spare parts strategy and inventory planning with supply chain and site teams.
  • Partner with Hardware Engineering, Manufacturing, and Infrastructure teams to provide field feedback that improves future platform designs.
  • Develop scalable operational standards and best practices that can be deployed across future Stargate campuses.
  • Mentor on-site technicians and partner teams on advanced troubleshooting methodologies and hardware operational excellence.

Skills

Datacenter hardware
GPU systems
Troubleshooting
Root cause analysis
Cross-functional leadership
Hardware reliability

Tools

FRACAS
RCCA
5-Why
Fishbone
FMEA
Telemetry platforms
Hardware monitoring tools

Job description

OpenAI seeks a Hardware Operations Engineer to serve as the technical authority for hardware reliability and fleet health at a flagship AI campus. You will operate at the intersection of hardware operations, sustaining engineering, and fleet reliability, partnering with CSPs, fleet-health engineers, hardware teams, and OEMs to diagnose and resolve production hardware issues.

The role drives RCA investigations, reliability improvement initiatives, lifecycle programs, and operational readiness,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Datacenter Hardware Operations Lead
Senior Datacenter Hardware Operations Lead

OpenAI • United States

On-site
USD 86,400 - 228,000
Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • Seattle (WA)

On-site
USD 150,000 - 200,000
Lead Datacenter Hardware Technician
Lead Datacenter Hardware Technician

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead Data Center Hardware Reliability Engineer
Lead Data Center Hardware Reliability Engineer

OpenAI • United States

Remote
USD 150,000 - 210,000
Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • United States

On-site
USD 86,400 - 228,000
Senior Data Center Hardware Reliability Engineer
Senior Data Center Hardware Reliability Engineer

Gohyred • San Francisco (CA)

On-site
USD 180,000 - 280,000
Technical Lead, AI Hardware Deployment & Operations
Technical Lead, AI Hardware Deployment & Operations

OpenAI • United States

Remote
USD 180,000 - 240,000
Senior Data Center Hardware Reliability Engineer
Senior Data Center Hardware Reliability Engineer

OpenAI, Inc. • San Francisco (CA)

On-site
USD 226,000 - 285,000
Relocation support
Daily meals in our offices
Learning and development stipend
+1
Hardware-Focused Site Reliability Engineer – Data Center
Hardware-Focused Site Reliability Engineer – Data Center

Maven Ventures • Southaven (MS)

On-site
USD 90,000 - 130,000