Mach aus dieser Rolle ein Bewerbungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.
Luma is building unified general intelligence capable of operating in the physical world. We seek an experienced RL at scale engineer to design, build, and scale distributed RL post‑training systems that coordinate trainer, rollout, environment, and reward workloads across thousands of GPUs.
You will own high‑throughput rollout generation, integrate inference engines like vLLM, implement weight synchronization, and support asynchronous/off‑policy schemes while developing agentic environments for
You'll build the systems that make reinforcement learning work at frontier scale — coupling policy optimization with large fleets of inference workers, agentic environments, and the reward and verification systems that turn model behavior into learning signal. RL is how Luma's models go from capable to useful.
RL at scale is a full-loop systems problem: training, rollout generation, environment execution, and reward computation running concurrently across thousands of GPUs, all needing to stay fast, stable, and correct together. It fits someone who has lived this — post-trained LLMs with RL, built environments and verifiers, and debugged asynchronous rollout pipelines at scale. If you haven't operated RL at real scale, this will be deep water.
One way the first 90 could unfold.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
Compensation Range: $195K - $395K