Software Engineer, Hardware Health

OpenAI

San Francisco (CA)

On-site

USD 250,000 - 445,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a Software Engineer for Hardware Health in San Francisco. The role involves maintaining the health of compute clusters, building automated systems for monitoring hardware, and ensuring efficient operations across large-scale distributed environments. Candidates should have over 7 years of experience in software engineering, strong skills in Python and shell scripting, and an interest in infrastructure platforms. Compensation ranges from $250K to $445K plus equity, reflecting the demanding nature of the work.

Qualifications

  • 7+ years of industry experience in software or infrastructure engineering.
  • Strong proficiency with Python and shell scripting.
  • Experience building large-scale distributed systems or infrastructure platforms.

Responsibilities

  • Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
  • Build automation and tooling that enables global cluster management with minimal manual intervention.
  • Investigate hardware failures and system-level issues across large-scale compute environments.

Skills

Python
Shell scripting
SQL
PromQL

Job description

Software Engineer, Hardware Health

Frontiers Clusters - San Francisco

About the Team

The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet.

Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling.

We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads.

About the Role

On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale.

Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams.

Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally.

Responsibilities
  • Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
  • Build and evolve health checks that detect, remediate, and verify failures at scale.
  • Ensure critical health checks execute with minimal latency to maximize workload uptime.
  • Investigate hardware failures and system‑level issues across large-scale compute environments.
  • Own node lifecycle workflows including drain, quarantine, repair, RMA, and return‑to‑service processes.
  • Build automation and tooling that enables global cluster management with minimal manual intervention.
  • Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.
Qualifications
  • 7+ years of industry experience in software or infrastructure engineering.
  • Strong proficiency with Python and shell scripting.
  • Experience building large‑scale distributed systems or infrastructure platforms.
  • Comfort digging into noisy operational data using SQL, PromQL, or similar tooling.
  • Experience building reproducible analyses and operational tooling.
  • Strong systems debugging and operational instincts with an ownership mindset.
Bonus if you have
  • Experience with low‑level hardware systems and Linux tooling (e.g., PCIe, InfiniBand, RoCE, networking, power management, kernel performance tuning, FW/SW debugging).
  • Experience operating or debugging large‑scale GPU or accelerator clusters.
  • Expertise in network operations, observability, or systems telemetry.
  • Experience with automated remediation systems or fleet lifecycle management.
  • Experience improving reliability, utilization, or workload uptime in distributed compute environments.
About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general‑purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for U.S.‑based candidates. For unincorporated Los Angeles County workers, criminal history may be considered in relation to duties such as protecting computer hardware, returning all hardware upon termination, and safeguarding proprietary information.

Compensation

$250K – $445K + equity

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • United States

On-site
USD 86,000 - 228,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Software Engineer, Hardware
Software Engineer, Hardware

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Hardware Health
Software Engineer, Hardware Health

Slope • San Francisco (CA)

On-site
USD 130,000 - 160,000
Software Engineer, Infrastructure, Consumer Devices
Software Engineer, Infrastructure, Consumer Devices

OpenAI • Los Angeles (CA)

Hybrid
USD 325,000 - 440,000
Tech Lead, Deployment & Operations - Custom Infrastructure
Tech Lead, Deployment & Operations - Custom Infrastructure

OpenAI • San Francisco (CA)

On-site
USD 342,000 - 445,000
Software Engineer - Data Aquisition (systems)
Software Engineer - Data Aquisition (systems)

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 405,000
Equity
Software Engineer, Compute Foundations Systems
Software Engineer, Compute Foundations Systems

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000