Distributed ML Training Performance Engineer

OpenAI

California (MO)

Hybrid

USD 170,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance

Job summary

OpenAI is seeking a Training Performance Engineer in San Francisco, CA, to push throughput and uptime across our distributed training stack. You will analyze large-scale runs, pinpoint bottlenecks, and design practical optimizations that scale with growing model size while maintaining compute efficiency.

You will work closely with runtime and systems engineers, researchers, and product teams to implement kernel optimizations, scheduling improvements, and data movement strategies.

Qualifications

  • Experience analyzing and optimizing distributed training workloads.
  • Strong programming skills in Python and C++.
  • Familiarity with PyTorch, JAX, or TensorFlow.
  • Experience running multi-GPU systems or HPC clusters.
  • Ability to translate profiling data into engineering improvements.

Responsibilities

  • Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage.
  • Optimize GPU utilization and throughput for large-scale distributed model training.
  • Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
  • Implement model graph transforms to improve end to end throughput.
  • Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.
  • Partner with researchers to ensure new model architectures scale efficiently during pre-training.
  • Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs.

Skills

Python
C++
Rust
CUDA

Tools

NCCL
MPI
UCX
PyTorch
JAX
TensorFlow

Job description

OpenAI is seeking a Training Performance Engineer in San Francisco, CA, to push throughput and uptime across our distributed training stack. You will analyze large-scale runs, pinpoint bottlenecks, and design practical optimizations that scale with growing model size while maintaining compute efficiency.

You will work closely with runtime and systems engineers, researchers, and product teams to implement kernel optimizations, scheduling improvements, and data movement strategies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Training Performance Engineer
Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance
Training Performance Engineer
Training Performance Engineer

Slope • San Francisco (CA)

On-site
USD 250,000 - 460,000
Relocation assistance
Flexible working hours
Collaborative work environment
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Training: ML Framework Engineer
Training: ML Framework Engineer

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Relocation assistance
Hybrid work model
ML Framework Engineer: Accelerate Large-Scale Training
ML Framework Engineer: Accelerate Large-Scale Training

OpenAI • California (MO)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work model
OpenAI team culture
Training: ML Framework Engineer
Training: ML Framework Engineer

OpenAI • California (MO)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work model
OpenAI team culture
Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
AI Training Performance Engineer — Large-Scale GPU Training
AI Training Performance Engineer — Large-Scale GPU Training

Figure • San Jose (CA)

On-site
USD 200,000 - 400,000
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Applied ML Engineer: Scale & Deploy Large Models
Applied ML Engineer: Scale & Deploy Large Models

AI Breaking Wire • San Francisco (CA)

Hybrid
USD 250,000 - 380,000
Equity
Medical, dental, and vision
Unlimited PTO
+2