GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI

New York (NY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Hybrid work model

Job summary

OpenAI in San Francisco invites engineers to join the Fleet Infrastructure team, shaping the world's largest GPU fleet for model training and deployment. The role involves designing, deploying and operating infrastructure systems for scalable compute and research workloads.

You will collaborate with researchers and product teams, implement scheduling and provisioning, and drive high utilization with reliable services; relocation support offered and a hybrid work model with three days in the

Qualifications

  • Experience with hyperscale compute systems.
  • Strong programming skills and ability to interface with researchers and product teams.
  • Familiarity with public clouds (Azure) and Kubernetes.

Responsibilities

  • Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
  • Interface with researchers and product teams to understand workload requirements.
  • Collaborate with hardware, infrastructure, and business teams to provide a high-utilization, reliable service.

Skills

Hyperscale compute
Programming skills
Public clouds (Azure)
Kubernetes
Execution-focused mentality
AI/ML workloads

Job description

OpenAI in San Francisco invites engineers to join the Fleet Infrastructure team, shaping the world's largest GPU fleet for model training and deployment. The role involves designing, deploying and operating infrastructure systems for scalable compute and research workloads.

You will collaborate with researchers and product teams, implement scheduling and provisioning, and drive high utilization with reliable services; relocation support offered and a hybrid work model with three days in the

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

OpenAI • New York (NY)

On-site
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
GPU Fleet Orchestrator for AI Infra
GPU Fleet Orchestrator for AI Infra

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 210,000
AI Infrastructure Engineer — Automate GPU Fleets (Hybrid)
AI Infrastructure Engineer — Automate GPU Fleets (Hybrid)

WeHireYou • New Amsterdam (IN)

Hybrid
USD 113,000 - 181,000
Principal AI Infrastructure Engineer — GPU Fleet
Principal AI Infrastructure Engineer — GPU Fleet

Nscale • Bellevue (WA)

On-site
USD 240,000 - 400,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
GPU Fleet Orchestrator for AI Infrastructure
GPU Fleet Orchestrator for AI Infrastructure

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

Cloudjobs • New York (NY)

On-site
USD 140,000 - 210,000
Relocation assistance
Compute Foundations Engineer — Kubernetes & GPU Infra
Compute Foundations Engineer — Kubernetes & GPU Infra

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 490,000