AI Infrastructure Engineer, pAGI

OpenAI

United States

Remote

USD 150,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OpenAI is seeking an AI Systems Engineer to scale infrastructure for training and evaluation workflows. You will own projects from bottleneck analysis to deployment, merging distributed systems engineering with performance optimization, close collaboration with researchers, and platform development.

You will build and operate infrastructure for large-scale training and evaluation, develop shared inference and grading platforms with health monitoring, and improve scheduling to maximize GPU

Qualifications

  • Experience building or operating large-scale distributed systems.
  • Strong software engineering fundamentals.
  • Ability to design, deploy, and operate infrastructure for ML workloads.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures.
  • Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance.
  • Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention.

Skills

Distributed systems engineering
Performance optimization
Software engineering fundamentals
Research tooling

Job description

About the Team

pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model.

About the Role

We're looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You'll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production - directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer, pAGI
AI Infrastructure Engineer, pAGI

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
AI Infrastructure Engineer, pAGI
AI Infrastructure Engineer, pAGI

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Infrastructure Engineer — Scale Training & Inference
AI Infrastructure Engineer — Scale Training & Inference

OpenAI • United States

Remote
USD 150,000 - 190,000
AI Infrastructure Engineer, pAGI
AI Infrastructure Engineer, pAGI

Triwill Group • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 190,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
AI Engineer: Model Training, Inference & GPU Infra
AI Engineer: Model Training, Inference & GPU Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000