Stand out for this role — generate a tailored resume and cover letter in about a minute.
Stealth AI Lab in London or Paris (Hybrid) is seeking a Member of Technical Staff for post-training RL work. You’ll improve models after training and enhance the systems that run experiments at scale.
You’ll explore RL methods like PPO, GRPO and SDPO, own training loops, and boost throughput, GPU utilization, and memory efficiency while debugging training vs inference differences.
Stealth AI Lab | London or Paris (Hybrid)
This role is about making models better after their initial training. You'll work on reinforcement learning and other post-training methods, while also improving the systems needed to run those experiments efficiently at scale.
The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.
You'll work across both the learning algorithms and the infrastructure underneath them. That means going from an RL experiment to a GPU profiler trace, finding what's limiting performance, and making sure systems improvements don't change the way the model learns.
Shortlisted candidates will be contacted within 48 hours.