Lead RL Infra Engineer: Scale Training & Environments

Luma

Greater London

Presencial

GBP 147.000 - 299.000

Jornada completa

hace 9 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

Luma is building unified general intelligence capable of operating in the physical world. We seek an experienced RL at scale engineer to design, build, and scale distributed RL post‑training systems that coordinate trainer, rollout, environment, and reward workloads across thousands of GPUs.

You will own high‑throughput rollout generation, integrate inference engines like vLLM, implement weight synchronization, and support asynchronous/off‑policy schemes while developing agentic environments for

Formación

  • Hands-on experience with post‑training LLMs using RL at meaningful scale.
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents.

Responsabilidades

  • Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
  • Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.
  • Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction.
  • Build reward infrastructure: verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
  • Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable.
  • Advance training efficiency and stability, and turn new post-training ideas into production runs with researchers.

Conocimientos

RL at scale
PPO/GRPO RL
PyTorch distributed training
vLLM integration
SGLang
NCCL/MPI networking
RLHF experience
GPU clusters

Herramientas

PyTorch
vLLM
SGLang
Ray
OpenRLHF

Descripción del empleo

Luma is building unified general intelligence capable of operating in the physical world. We seek an experienced RL at scale engineer to design, build, and scale distributed RL post‑training systems that coordinate trainer, rollout, environment, and reward workloads across thousands of GPUs.

You will own high‑throughput rollout generation, integrate inference engines like vLLM, implement weight synchronization, and support asynchronous/off‑policy schemes while developing agentic environments for

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Híbrido
GBP 120.000 - 180.000
RL Systems Engineer: Scalable Post-Training Infrastructure
RL Systems Engineer: Scalable Post-Training Infrastructure

Luma • Gran Bretaña

Presencial
GBP 150.000 - 190.000
Lead RL Infrastructure Engineer - Scale Training Systems
Lead RL Infrastructure Engineer - Scale Training Systems

Speedrun Talent Network • Greater London

Híbrido
GBP 120.000 - 190.000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

AItoolnavio • Greater London

Híbrido
GBP 120.000 - 180.000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Speedrun Talent Network • Greater London

Híbrido
GBP 120.000 - 190.000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Greater London

Presencial
GBP 147.000 - 299.000
Tech Lead Manager (MLRE, ML Systems)
Tech Lead Manager (MLRE, ML Systems)

Scale AI • York and North Yorkshire

Presencial
GBP 120.000 - 180.000
Health & Wellbeing
Career Growth stipend
Community events
+1
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • Gran Bretaña

Presencial
GBP 90.000 - 150.000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Híbrido
GBP 120.000 - 180.000
Inference Systems Engineer (GPU/Cluster Scaling)
Inference Systems Engineer (GPU/Cluster Scaling)

AItoolnavio • Greater London

Híbrido
GBP 100.000 - 180.000