GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI

New York (NY)

Hybrid

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance
Hybrid work model

Job summary

OpenAI in San Francisco invites engineers to join the Fleet Infrastructure team, shaping the world's largest GPU fleet for model training and deployment. The role involves designing, deploying and operating infrastructure systems for scalable compute and research workloads.

You will collaborate with researchers and product teams, implement scheduling and provisioning, and drive high utilization with reliable services; relocation support offered and a hybrid work model with three days in the

Qualifications

  • Experience with hyperscale compute systems.
  • Strong programming skills and ability to interface with researchers and product teams.
  • Familiarity with public clouds (Azure) and Kubernetes.

Responsibilities

  • Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
  • Interface with researchers and product teams to understand workload requirements.
  • Collaborate with hardware, infrastructure, and business teams to provide a high-utilization, reliable service.

Skills

Hyperscale compute
Programming skills
Public clouds (Azure)
Kubernetes
Execution-focused mentality
AI/ML workloads

Job description

OpenAI in San Francisco invites engineers to join the Fleet Infrastructure team, shaping the world's largest GPU fleet for model training and deployment. The role involves designing, deploying and operating infrastructure systems for scalable compute and research workloads.

You will collaborate with researchers and product teams, implement scheduling and provisioning, and drive high utilization with reliable services; relocation support offered and a hybrid work model with three days in the

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
GPU Fleet Engineering Manager – AI Infra Leader
GPU Fleet Engineering Manager – AI Infra Leader

Lambda • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage foryou
Wellness and commuter stipends
401k with 2% company match (USA)
+1
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
Software Engineer, Fleet Management
Software Engineer, Fleet Management

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Software Engineer, Fleet Management
Software Engineer, Fleet Management

OpenAI • San Francisco (CA)

Hybrid
USD 230,000 - 490,000