Distributed RL Systems Engineer — Scale Training & Inference

Luma AI

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma AI is building scalable reinforcement learning systems that couple policy optimization with thousands of GPUs, dispatching trainer, rollout, environment, and reward workloads. You will design, build, and scale these post-training systems to run at frontier scale.

The role involves creating high-throughput rollout generation, integrating inference engines like vLLM and SGLang, and developing robust reward and evaluation tooling to keep experiments fast, stable, and correct across large

Qualifications

  • Post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at scale.
  • Extensive distributed PyTorch training and parallelism for foundation models.
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use.
  • Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.

Responsibilities

  • Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
  • Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.
  • Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction — reproducible and scalable to millions of episodes.
  • Build reward infrastructure: verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
  • Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable.
  • Advance training efficiency and stability, and turn new post-training ideas into production runs with researchers.

Skills

Post-training RL with LLMs
Distributed PyTorch training
RL environments & verifiers
Rollout inference engines
GPU cluster management
NCCL/MPI networking

Tools

veRL
OpenRLHF
TRL
Ray orchestration
vLLM
SGLang

Job description

Luma AI is building scalable reinforcement learning systems that couple policy optimization with thousands of GPUs, dispatching trainer, rollout, environment, and reward workloads. You will design, build, and scale these post-training systems to run at frontier scale.

The role involves creating high-throughput rollout generation, integrating inference engines like vLLM and SGLang, and developing robust reward and evaluation tooling to keep experiments fast, stable, and correct across large

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
ML Inference Systems Engineer — Kubernetes & GPU Scale
ML Inference Systems Engineer — Kubernetes & GPU Scale

Luma AI • United States

Remote
USD 140,000 - 190,000
RL Infrastructure Engineer: Scale End-to-End ML
RL Infrastructure Engineer: Scale End-to-End ML

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Medical benefits
Dental benefits
Vision benefits
+7
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior RL Post-Training Systems Engineer (Equity)
Senior RL Post-Training Systems Engineer (Equity)

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000