RL Training Infra Engineer: Scale, Debug, & Optimize

OpenAI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is looking for a strong generalist engineer to focus on reinforcement learning training runs. The role requires debugging across various infrastructures and improving reliability and efficiency in large-scale systems.

The ideal candidate is highly motivated, can work in messy areas, and enjoys building systems that support cutting-edge AI technologies. Applicants should have experience in ML infrastructure, debugging skills, and a strong attention to detail.

Qualifications

  • Experience in reinforcement learning, inference, and scaling.
  • Ability to debug across training systems and infrastructures.
  • Strong communication skills and high ownership.

Responsibilities

  • Keep large-scale RL training runs moving by solving urgent problems.
  • Debug issues in training systems and distributed infrastructure.
  • Improve the reliability and efficiency of RL training runs.

Skills

Generalist engineering
Debugging skills
Experience in ML infrastructure
High attention to detail
Ability to solve technical problems

Job description

OpenAI is looking for a strong generalist engineer to focus on reinforcement learning training runs. The role requires debugging across various infrastructures and improving reliability and efficiency in large-scale systems.

The ideal candidate is highly motivated, can work in messy areas, and enjoys building systems that support cutting-edge AI technologies. Applicants should have experience in ML infrastructure, debugging skills, and a strong attention to detail.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
RL Infrastructure Engineer — Scale & Production
RL Infrastructure Engineer — Scale & Production

Inception • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, dental, and vision insurance
Catered meals (breakfast, lunch, &
Equity
+1
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Systems Engineer: Training Infra
Senior ML Systems Engineer: Training Infra

Neura Market • San Francisco (CA)

On-site
USD 295,000 - 380,000
Relocation assistance
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000