RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio

Greater London

Hybrid

GBP 120.000 - 180.000

Vollzeit

Vor 2 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Schick keinen Standard-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Luma is building unified general intelligence that can generate, understand, and operate in the physical world. We seek a senior RL systems engineer to design, build, and scale post-training RL workloads across thousands of GPUs, orchestrating trainer, rollout, environment, and reward components for production-scale runs.

You will own end-to-end tooling for reliable, fast RL loops and collaborate with researchers to turn ideas into robust, scalable implementations.

Qualifikationen

  • Hands-on experience with post-training LLM RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale.
  • Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert) for foundation models.
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents.
  • Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.

Aufgaben

  • Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
  • Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.
  • Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction — reproducible and scalable to millions of episodes.
  • Build reward infrastructure: verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
  • Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable.
  • Advance training efficiency and stability, and turn new post-training ideas into production runs with researchers.

Kenntnisse

RL post-training
Distributed PyTorch
RL environments
RL frameworks
GPU networking

Tools

veRL
OpenRLHF
TRL
Ray orchestration
vLLM
SGLang

Jobbeschreibung

Luma is building unified general intelligence that can generate, understand, and operate in the physical world. We seek a senior RL systems engineer to design, build, and scale post-training RL workloads across thousands of GPUs, orchestrating trainer, rollout, environment, and reward components for production-scale runs.

You will own end-to-end tooling for reliable, fast RL loops and collaborate with researchers to turn ideas into robust, scalable implementations.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Lead RL Infra Engineer: Scale Training & Environments
Lead RL Infra Engineer: Scale Training & Environments

Luma • Greater London

Vor Ort
GBP 147.000 - 299.000
RL Systems Engineer: Scalable Post-Training Infrastructure
RL Systems Engineer: Scalable Post-Training Infrastructure

Luma • Großbritannien

Vor Ort
GBP 150.000 - 190.000
Lead RL Infrastructure Engineer - Scale Training Systems
Lead RL Infrastructure Engineer - Scale Training Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120.000 - 190.000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

AItoolnavio • Greater London

Hybrid
GBP 120.000 - 180.000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Speedrun Talent Network • Greater London

Hybrid
GBP 120.000 - 190.000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Greater London

Vor Ort
GBP 147.000 - 299.000
Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

Vor Ort
GBP 147.000 - 299.000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120.000 - 180.000
Senior Software Engineer, Large-Scale Model Inference
Senior Software Engineer, Large-Scale Model Inference

Luma • Greater London

Vor Ort
GBP 147.000 - 261.000
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • Großbritannien

Vor Ort
GBP 90.000 - 150.000