RL Systems Engineer: Scalable Post-Training Infrastructure

Luma

United Kingdom

On-site

GBP 150,000 - 190,000

Full time

30 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma in the United Kingdom is building RL systems at frontier scale, coupling policy optimization with thousands of GPUs. You will design, build, and scale post-training pipelines that coordinate trainer, rollout, environments, and reward workloads.

We seek someone with hands-on experience training LLMs with RL, expertise in distributed PyTorch, and familiarity with inference engines like vLLM and SGLang. You’ll help harden the loop, improve throughput, and ship robust evaluation tools.

Qualifications

  • Hands-on experience with post‑training LLMs using RL (PPO/GRPO family, RLHF, RLVR).
  • Extensive distributed PyTorch training and model parallelism for foundation models.
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents.
  • Familiarity with RL post‑training frameworks (veRL, OpenRLHF, TRL) and rollout inference engines.
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI).

Responsibilities

  • Design, build, and scale distributed RL post‑training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
  • Build high‑throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off‑policy schemes.
  • Design RL environments for agentic, multi‑step tasks — sandboxed code execution, tool use, computer use, multimodal interaction.
  • Build reward infrastructure: verifiable/programmatic rewards, reward‑model serving, LLM‑as‑judge pipelines, and defenses against reward hacking.
  • Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable.
  • Advance training efficiency and stability, and turn new post‑training ideas into production runs with researchers.
  • Days 1–30 — Immerse & Diagnose: Learn the current RL stack and where throughput, stability, or correctness break.
  • Days 30–60 — Ship & Validate: Improve a piece of the loop and prove it on a real run.
  • Days 60–90 — Scale & Systemize: Harden the full loop across thousands of GPUs and asynchronous architectures.

Skills

Post-training RL
Distributed PyTorch
RL environments
GPGPU systems
RLHF / PPO

Tools

vLLM
SGLang
veRL
OpenRLHF
TRL
Ray
Kubernetes
NCCL
MPI
PyTorch
CUDA

Job description

Luma in the United Kingdom is building RL systems at frontier scale, coupling policy optimization with thousands of GPUs. You will design, build, and scale post-training pipelines that coordinate trainer, rollout, environments, and reward workloads.

We seek someone with hands-on experience training LLMs with RL, expertise in distributed PyTorch, and familiarity with inference engines like vLLM and SGLang. You’ll help harden the loop, improve throughput, and ship robust evaluation tools.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000
Lead RL Infra Engineer: Scale Training & Environments
Lead RL Infra Engineer: Scale Training & Environments

Luma • Greater London

On-site
GBP 147,000 - 299,000
Lead RL Infrastructure Engineer - Scale Training Systems
Lead RL Infrastructure Engineer - Scale Training Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 190,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 190,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Greater London

On-site
GBP 147,000 - 299,000
Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

On-site
GBP 147,000 - 299,000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Tech Lead Manager (MLRE, ML Systems)
Tech Lead Manager (MLRE, ML Systems)

Scale AI • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Health & Wellbeing
Career Growth stipend
Community events
+1
Research Engineer (Post-Training)
Research Engineer (Post-Training)

Axiōma Search • Greater London

Hybrid
GBP 90,000 - 140,000