AI Systems Engineer - Scalable Training Infra

OpenAI

San Francisco (CA)

On-site

USD 160,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

OpenAI is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck analysis to deployment, combining distributed systems engineering with performance optimization and close collaboration with researchers.

You might build shared grading services, improve resource allocation, or bring a new training stack into production—directly accelerating research progress and reliability across OpenAI's pipelines.

Qualifications

  • Strong software engineering fundamentals and experience building or operating large-scale distributed systems.
  • Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
  • Highly self-motivated and comfortable taking ownership of open-ended problems.
  • Enjoy debugging across system boundaries and using measurements to guide improvements in performance and reliability.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures.
  • Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance.
  • Build self-service tools, automated validation, and observability to help researchers launch experiments and compare results with less manual intervention.

Skills

Distributed systems
Software engineering
ML infrastructure
Performance optimization
Ownership of open-ended problems

Job description

OpenAI is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck analysis to deployment, combining distributed systems engineering with performance optimization and close collaboration with researchers.

You might build shared grading services, improve resource allocation, or bring a new training stack into production—directly accelerating research progress and reliability across OpenAI's pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Systems Engineer - Scale Training & Inference
AI Systems Engineer - Scale Training & Inference

United States Digital Space LLC • United States

Remote
USD 140,000 - 190,000
Performance Engineer, Scalable AI Systems
Performance Engineer, Scalable AI Systems

Coinscapture • Northern (KY)

Hybrid
USD 140,000 - 230,000
Senior Systems Engineer for Scalable AI Training Infra
Senior Systems Engineer for Scalable AI Training Infra

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
RL Training Infra Engineer: Scale, Debug, & Optimize
RL Training Infra Engineer: Scale, Debug, & Optimize

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Infra Engineer for Scalable Enterprise Platforms
Senior AI Infra Engineer for Scalable Enterprise Platforms

Seekr • Austin (TX)

Hybrid
USD 180,000 - 260,000
Equity Ownership
Unlimited PTO and holidays
Hybrid work environment (Reston, VA &\
+2