GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI

New York (NY)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

OpenAI is seeking a Software Engineer for the Fleet High Performance Computing (HPC) team in New York to ensure reliability and uptime for OpenAI's compute fleet and data centers. You will build automation, monitor health, and optimize lifecycle events across clusters and hardware.

You will collaborate with networking and infra teams, develop tooling in Python/Go, and push for scalable solutions that reduce manual work while improving performance and reliability in a demanding HPC environment.

Qualifications

  • Experience building and operating large-scale server environments.
  • Strong Python/Go skills and ability to diagnose complex systems.
  • Solid knowledge of Linux, networking, and hardware fundamentals.

Responsibilities

  • Build and maintain automation systems for provisioning and managing server fleets.
  • Develop tools to monitor server health, performance, and lifecycle events.
  • Collaborate with clusters, networking, and infrastructure teams.
  • Partner with external operators to ensure a high level of quality.
  • Identify and fix performance bottlenecks and inefficiencies.
  • Continuously improve automation to reduce manual work.

Skills

Python
Go
Linux
Networking
SQL
PromQL
Pandas

Tools

Prometheus
Grafana

Job description

OpenAI is seeking a Software Engineer for the Fleet High Performance Computing (HPC) team in New York to ensure reliability and uptime for OpenAI's compute fleet and data centers. You will build automation, monitor health, and optimize lifecycle events across clusters and hardware.

You will collaborate with networking and infra teams, develop tooling in Python/Go, and push for scalable solutions that reduce manual work while improving performance and reliability in a demanding HPC environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

Cloudjobs • New York (NY)

On-site
USD 120,000 - 160,000
AI Infrastructure & Automation Engineer
AI Infrastructure & Automation Engineer

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Fleet Automation Engineer - HPC Infrastructure
Fleet Automation Engineer - HPC Infrastructure

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000
Lunch stipend
Medical benefits
Paid parental leave
+3
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

OpenAI • New York (NY)

On-site
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Infra Software Engineer: Fleet & HPC Automation
Infra Software Engineer: Fleet & HPC Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Highly competitive package with equity
Flexible workplace
Medical, dental, vision benefits
+2
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000