Distributed RL Systems Engineer — Scale Training & Inference

Luma AI

United States

Remote

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma AI is building scalable reinforcement learning systems that couple policy optimization with thousands of GPUs, dispatching trainer, rollout, environment, and reward workloads. You will design, build, and scale these post-training systems to run at frontier scale.

The role involves creating high-throughput rollout generation, integrating inference engines like vLLM and SGLang, and developing robust reward and evaluation tooling to keep experiments fast, stable, and correct across large

Qualifications

  • Post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at scale.
  • Extensive distributed PyTorch training and parallelism for foundation models.
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use.
  • Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.

Responsibilities

  • Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
  • Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.
  • Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction — reproducible and scalable to millions of episodes.
  • Build reward infrastructure: verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
  • Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable.
  • Advance training efficiency and stability, and turn new post-training ideas into production runs with researchers.

Skills

Post-training RL with LLMs
Distributed PyTorch training
RL environments & verifiers
Rollout inference engines
GPU cluster management
NCCL/MPI networking

Tools

veRL
OpenRLHF
TRL
Ray orchestration
vLLM
SGLang

Job description

Luma AI is building scalable reinforcement learning systems that couple policy optimization with thousands of GPUs, dispatching trainer, rollout, environment, and reward workloads. You will design, build, and scale these post-training systems to run at frontier scale.

The role involves creating high-throughput rollout generation, integrating inference engines like vLLM and SGLang, and developing robust reward and evaluation tooling to keep experiments fast, stable, and correct across large

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000