GPU HPC Systems Engineer — Fleet Reliability & Automation

OpenAI

California (MO)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI's Fleet HPC team seeks a software engineer to ensure reliability and uptime of our compute fleet. You will build automation for provisioning, monitoring, and lifecycle events across large-scale server environments, collaborating with clusters, networking, and infrastructure teams.

You will own system-level investigations, develop tooling in Python/Go, analyze noisy data with SQL/PromQL, and push for automated remediation at scale while maintaining high safety and efficiency in our

Qualifications

  • Experience managing large-scale server environments.
  • A balance of strengths in building and operationalizing.
  • Proficiency in Python, Go, or similar languages.
  • Strong Linux, networking, and server hardware knowledge.
  • Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.

Responsibilities

  • Build and maintain automation systems for provisioning and managing server fleets.
  • Develop tools to monitor server health, performance, and lifecycle events.
  • Collaborate with clusters, networking, and infrastructure teams.
  • Partner with external operators to ensure a high level of quality.
  • Identify and fix performance bottlenecks and inefficiencies.
  • Continuously improve automation to reduce manual work.

Skills

Python
Go
Linux
Networking
Server hardware
SQL/PromQL

Tools

Prometheus
Grafana
IPMI

Job description

OpenAI's Fleet HPC team seeks a software engineer to ensure reliability and uptime of our compute fleet. You will build automation for provisioning, monitoring, and lifecycle events across large-scale server environments, collaborating with clusters, networking, and infrastructure teams.

You will own system-level investigations, develop tooling in Python/Go, analyze noisy data with SQL/PromQL, and push for automated remediation at scale while maintaining high safety and efficiency in our

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fleet Hardware Reliability Engineer (Automation & HPC)
Fleet Hardware Reliability Engineer (Automation & HPC)

OpenAI • California (MO)

On-site
USD 180,000 - 270,000
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • California (MO)

On-site
USD 180,000 - 260,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Fleet Automation Engineer — HPC Infrastructure
Fleet Automation Engineer — HPC Infrastructure

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 160,000
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Fleet Automation Engineer – HPC & GPU Compute
Fleet Automation Engineer – HPC & GPU Compute

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 180,000
Company-Paid Lunch Stipend
Company-Paid Benefits: Medical, Dental
401(k) matching up to 6%
+2
Fleet & Automation Infrastructure Engineer — AI/HPC
Fleet & Automation Infrastructure Engineer — AI/HPC

Nscale • Houston (TX)

On-site
USD 150,000 - 215,000
Base salary + equity
Annual reviews
Growth opportunities
GPU Fleet Infra Engineer — Scale & Automate
GPU Fleet Infra Engineer — Scale & Automate

OpenAI • California (MO)

Hybrid
USD 180,000 - 260,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000