AI Infrastructure Engineer: Scale Training & Systems

Precision Labs

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OpenAI is seeking an AI Systems Engineer to scale the infrastructure behind our training and evaluation workflows. You’ll own projects from bottleneck identification to deployment, working across distributed systems, performance optimization, and collaboration with researchers.

You’ll build shared platforms, optimize scheduling and resource usage, and create observability tools to accelerate research while maintaining reliability at scale.

Qualifications

  • Experience building and operating large-scale distributed systems.
  • Strong software engineering fundamentals.
  • Experience in ML infrastructure, inference systems, or infrastructure tooling.
  • Ability to own troubleshooting from bottlenecks to deployment.
  • Comfort with cross-team collaboration.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time.
  • Diagnose bottlenecks across training, inference, and orchestration and coordinate across teams to improve end-to-end performance.
  • Build self-service tools and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention.

Skills

Distributed systems
Performance optimization
Large-scale infrastructure
Research tooling
Observability

Job description

OpenAI is seeking an AI Systems Engineer to scale the infrastructure behind our training and evaluation workflows. You’ll own projects from bottleneck identification to deployment, working across distributed systems, performance optimization, and collaboration with researchers.

You’ll build shared platforms, optimize scheduling and resource usage, and create observability tools to accelerate research while maintaining reliability at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
RL Training Infra Engineer: Scale, Debug, & Optimize
RL Training Infra Engineer: Scale, Debug, & Optimize

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
AI Training Infrastructure Engineer - Scale & Performance
AI Training Infrastructure Engineer - Scale & Performance

SupportFinity™ • San Francisco (CA)

On-site
USD 175,000 - 220,000
Meaningful equity
Competitive salary
Comprehensive benefits package
Platform Engineer - AI Systems & Cloud Infrastructure
Platform Engineer - AI Systems & Cloud Infrastructure

Scale AI • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
Health insurance
Dental & Vision
Retirement benefits
+3
Senior AI Infra Architect — Scale Enterprise AI
Senior AI Infra Architect — Scale Enterprise AI

Seekr • Washington

Hybrid
USD 170,000 - 260,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+2
Performance Engineer, Scalable AI Systems
Performance Engineer, Scalable AI Systems

Coinscapture • Northern (KY)

Hybrid
USD 140,000 - 230,000
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2