GPU Fleet Infra Engineer — Scale & Automate

OpenAI

California (MO)

Hybrid

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking an engineer for Fleet infrastructure to design, build, and operate scalable compute systems for model deployment and training on one of the world’s largest GPU fleets. You will work with researchers and product teams to understand workloads, develop scheduling, and automate cluster provisioning and upgrades.

The role emphasizes fast delivery, reliability, and close collaboration across hardware, infra, and software teams in a hybrid San Francisco setting.

Qualifications

  • Experience designing and operating infrastructure systems for large-scale compute.
  • Strong programming skills and ability to interface with researchers and product teams.

Responsibilities

  • Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
  • Interface with researchers and product teams to understand workload requirements.
  • Collaborate with hardware, infrastructure, and business teams to provide high utilization and reliability.

Skills

Programming
Hyperscale compute
Stakeholder communication

Tools

Kubernetes
Azure

Job description

OpenAI is seeking an engineer for Fleet infrastructure to design, build, and operate scalable compute systems for model deployment and training on one of the world’s largest GPU fleets. You will work with researchers and product teams to understand workloads, develop scheduling, and automate cluster provisioning and upgrades.

The role emphasizes fast delivery, reliability, and close collaboration across hardware, infra, and software teams in a hybrid San Francisco setting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

OpenAI • California (MO)

Hybrid
USD 180,000 - 260,000
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Software Engineer, OS & Fleet Orchestration at Scale
Software Engineer, OS & Fleet Orchestration at Scale

OpenAI • California (MO)

Hybrid
USD 170,000 - 250,000
Relocation assistance
Hybrid work model
AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
GPU HPC Systems Engineer — Fleet Reliability & Automation
GPU HPC Systems Engineer — Fleet Reliability & Automation

OpenAI • California (MO)

On-site
USD 180,000 - 260,000