Lead RL Infrastructure Engineer - Scale Training Systems

Speedrun Talent Network

Greater London

Hybrid

GBP 120,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Speedrun Talent Network seeks an expert to design, build, and scale distributed RL post-training systems for large-scale models. You will orchestrate trainer, rollout, environment, and reward workloads across thousands of GPUs and integrate inference engines such as vLLM and SGLang.

You will develop RL environments and reward infrastructures, verifiers, and evaluation harnesses, ensuring stable, high-throughput training and rollout across complex, asynchronous pipelines.

Qualifications

  • Hands-on experience with post-training LLMs using RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale.
  • Extensive distributed PyTorch training and parallelism for foundation models (FSDP, Tensor/Pipeline/Expert Parallel).
  • Experience building RL environments, reward functions, verifiers, and evaluation harnesses for LLM agents.

Responsibilities

  • Design, build, and scale distributed RL post-training systems across thousands of GPUs.
  • Develop high-throughput rollout generation and integrate inference engines (vLLM, SGLang).
  • Create sandboxed RL environments with multi-turn tool use and broad multimodal support.

Skills

Post-training LLMs with RL
Distributed PyTorch
RL environments & reward design
RL post-training frameworks
GPU clusters & NCCL/MPI

Tools

vLLM
SGLang
veRL
OpenRLHF
TRL
Ray

Job description

Speedrun Talent Network seeks an expert to design, build, and scale distributed RL post-training systems for large-scale models. You will orchestrate trainer, rollout, environment, and reward workloads across thousands of GPUs and integrate inference engines such as vLLM and SGLang.

You will develop RL environments and reward infrastructures, verifiers, and evaluation harnesses, ensuring stable, high-throughput training and rollout across complex, asynchronous pipelines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL Systems Engineer: Scalable Post-Training Infrastructure
RL Systems Engineer: Scalable Post-Training Infrastructure

Luma • United Kingdom

On-site
GBP 150,000 - 190,000
Lead RL Infra Engineer: Scale Training & Environments
Lead RL Infra Engineer: Scale Training & Environments

Luma • Greater London

On-site
GBP 147,000 - 299,000
RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 190,000
Tech Lead Manager (MLRE, ML Systems)
Tech Lead Manager (MLRE, ML Systems)

Scale AI • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Health & Wellbeing
Career Growth stipend
Community events
+1
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Greater London

On-site
GBP 147,000 - 299,000
Senior Research Scientist: Scale RL & Post-Training for Models
Senior Research Scientist: Scale RL & Post-Training for Models

PVH (Tommy Hilfiger/Calvin Klein) • Greater London

On-site
GBP 100,000 - 150,000
Member of Technical Staff — Environments / Evals
Member of Technical Staff — Environments / Evals

Axiōma Search • United Kingdom

Remote
GBP 90,000 - 130,000
Research Engineer (Post-Training)
Research Engineer (Post-Training)

Axiōma Search • Greater London

Hybrid
GBP 90,000 - 140,000