RL Systems Scientist — Multimodal Foundation Models

lumalabs-ai

San Francisco (CA)

On-site

USD 188,000 - 395,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma AI is recruiting engineers and scientists to design and operate the RL stack for our largest multimodal models. You will work on post-training RL, distributed training, and end-to-end reinforcement learning workflows across GPUs to drive capability and reliability.

You will collaborate with researchers to deploy scalable training, rollout, environments, reward systems, and evaluation tooling, balancing speed, stability and learning signal in production.

Qualifications

  • Hands-on experience post-training LLMs with reinforcement learning at meaningful scale.
  • Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models.
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use.
  • Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang).
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads.

Responsibilities

  • Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs
  • Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes
  • Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes
  • Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking
  • Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run
  • Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads
  • Collaborate closely with researchers to turn new post-training ideas (RLVR, agentic RL, long-horizon credit assignment, self-improvement loops) into production-quality training runs

Skills

Reinforcement learning
Distributed training
GPU systems
RL frameworks

Tools

vLLM
SGLang
PyTorch
Ray

Job description

Luma AI is recruiting engineers and scientists to design and operate the RL stack for our largest multimodal models. You will work on post-training RL, distributed training, and end-to-end reinforcement learning workflows across GPUs to drive capability and reliability.

You will collaborate with researchers to deploy scalable training, rollout, environments, reward systems, and evaluation tooling, balancing speed, stability and learning signal in production.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Multimodal AI Evaluation Engineer
Multimodal AI Evaluation Engineer

lumalabs-ai • New York (NY)

On-site
USD 190,000 - 375,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Research Scientist / Engineer — Multimodal Agent
Research Scientist / Engineer — Multimodal Agent

lumalabs-ai • San Francisco (CA)

On-site
USD 250,000 - 450,000
Senior Multimodal AI Evaluation Engineer
Senior Multimodal AI Evaluation Engineer

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000