AI Infrastructure Engineer — Scale Training & Inference

OpenAI

United States

Remote

USD 150,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OpenAI is seeking an AI Systems Engineer to scale infrastructure for training and evaluation workflows. You will own projects from bottleneck analysis to deployment, merging distributed systems engineering with performance optimization, close collaboration with researchers, and platform development.

You will build and operate infrastructure for large-scale training and evaluation, develop shared inference and grading platforms with health monitoring, and improve scheduling to maximize GPU

Qualifications

  • Experience building or operating large-scale distributed systems.
  • Strong software engineering fundamentals.
  • Ability to design, deploy, and operate infrastructure for ML workloads.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures.
  • Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance.
  • Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention.

Skills

Distributed systems engineering
Performance optimization
Software engineering fundamentals
Research tooling

Job description

OpenAI is seeking an AI Systems Engineer to scale infrastructure for training and evaluation workflows. You will own projects from bottleneck analysis to deployment, merging distributed systems engineering with performance optimization, close collaboration with researchers, and platform development.

You will build and operate infrastructure for large-scale training and evaluation, develop shared inference and grading platforms with health monitoring, and improve scheduling to maximize GPU

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI Infrastructure Engineer, pAGI
AI Infrastructure Engineer, pAGI

OpenAI • United States

Remote
USD 150,000 - 190,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Scaled Cognition • New York (NY)

On-site
USD 120,000 - 170,000
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
Software Engineer, AI Inference Infrastructure Platform
Software Engineer, AI Inference Infrastructure Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000