Turn this role into an interview — a resume and cover letter built around what this employer wants.
Luma AI is recruiting engineers and scientists to design and operate the RL stack for our largest multimodal models. You will work on post-training RL, distributed training, and end-to-end reinforcement learning workflows across GPUs to drive capability and reliability.
You will collaborate with researchers to deploy scalable training, rollout, environments, reward systems, and evaluation tooling, balancing speed, stability and learning signal in production.
Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
Reinforcement learning is how our foundation models go from capable to useful — learning to reason, use tools, and act over long horizons. The RL Infrastructure team builds the systems that make this possible at scale: high-throughput distributed training that couples policy optimization with large fleets of inference workers, environments that expose models to realistic multi-step tasks, and the reward, verification, and evaluation systems that turn model behavior into learning signal.
Unlike pretraining, RL at scale is a full-loop systems problem — training, rollout generation, environment execution, and reward computation all run concurrently across thousands of GPUs and must stay fast, stable, and correct together. We are looking for engineers and scientists who have lived this problem: people who have post-trained LLMs with RL, built environments and verifiers from scratch, and debugged what happens when an asynchronous rollout pipeline meets a frontier-scale training run. You will work alongside our research team to design and operate the RL stack for our largest multimodal models.
The base pay range for this role is $187,500 – $395,000 per year.