An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Luma in the United Kingdom is building RL systems at frontier scale, coupling policy optimization with thousands of GPUs. You will design, build, and scale post-training pipelines that coordinate trainer, rollout, environments, and reward workloads.
We seek someone with hands-on experience training LLMs with RL, expertise in distributed PyTorch, and familiarity with inference engines like vLLM and SGLang. You’ll help harden the loop, improve throughput, and ship robust evaluation tools.
Luma in the United Kingdom is building RL systems at frontier scale, coupling policy optimization with thousands of GPUs. You will design, build, and scale post-training pipelines that coordinate trainer, rollout, environments, and reward workloads.
We seek someone with hands-on experience training LLMs with RL, expertise in distributed PyTorch, and familiarity with inference engines like vLLM and SGLang. You’ll help harden the loop, improve throughput, and ship robust evaluation tools.