Fleet Hardware Reliability Engineer (Automation & HPC)

OpenAI

California (MO)

On-site

USD 180,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a software engineer for the Fleet Hardware team to ensure reliability and uptime across our compute fleet. You will build automation for provisioning, monitoring, and lifecycle events, collaborating with clusters, networking, and infra teams to minimize hardware failures that impact research and product services.

The role emphasizes deep problem solving, automation, and scalable health monitoring.

Qualifications

  • Experience managing large-scale server environments.
  • Proficiency in Python, Go, or similar languages.
  • Strong Linux, networking, and server hardware knowledge.
  • Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.

Responsibilities

  • Build and maintain automation systems for provisioning and managing server fleets.
  • Develop tools to monitor server health, performance, and lifecycle events.
  • Collaborate with clusters, networking, and infrastructure teams.
  • Partner with external operators to ensure a high level of quality.
  • Identify and fix performance bottlenecks and inefficiencies.
  • Continuously improve automation to reduce manual work.

Skills

Python
Go
Linux
Networking
SQL
PromQL
Pandas

Tools

Prometheus
Grafana

Job description

OpenAI is seeking a software engineer for the Fleet Hardware team to ensure reliability and uptime across our compute fleet. You will build automation for provisioning, monitoring, and lifecycle events, collaborating with clusters, networking, and infra teams to minimize hardware failures that impact research and product services.

The role emphasizes deep problem solving, automation, and scalable health monitoring.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU HPC Systems Engineer — Fleet Reliability & Automation
GPU HPC Systems Engineer — Fleet Reliability & Automation

OpenAI • California (MO)

On-site
USD 180,000 - 260,000
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, Fleet Hardware Health
Software Engineer, Fleet Hardware Health

OpenAI • California (MO)

On-site
USD 180,000 - 270,000
Datacenter Hardware Lead — Reliability & Fleet Health
Datacenter Hardware Lead — Reliability & Fleet Health

OpenAI • California (MO)

On-site
USD 120,000 - 180,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, Fleet Orchestration
Software Engineer, Fleet Orchestration

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Senior AI/ML Fleet Automation Engineer
Senior AI/ML Fleet Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
401(k) matching
Paid time off
Parental leave
+1
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • California (MO)

On-site
USD 180,000 - 260,000
AI Infrastructure Engineer - Fleet & Automation
AI Infrastructure Engineer - Fleet & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Equity
Base salary + equity
Career progression
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match