Research Engineer (Post-Training)

Axiōma Search

Greater London

Hybrid

GBP 90,000 - 140,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Stealth AI Lab in London or Paris (Hybrid) is seeking a Member of Technical Staff for post-training RL work. You’ll improve models after training and enhance the systems that run experiments at scale.

You’ll explore RL methods like PPO, GRPO and SDPO, own training loops, and boost throughput, GPU utilization, and memory efficiency while debugging training vs inference differences.

Qualifications

  • Strong programming and quantitative problem-solving skills.
  • Hands-on experience with PyTorch and model training.
  • Understanding of reinforcement learning or LLM post-training.
  • Experience with distributed training and GPU systems.
  • Strong experimental judgement.
  • Ability to work across ML research and systems engineering.

Responsibilities

  • Build and improve supervised fine-tuning, preference optimisation and RL methods
  • Work with approaches including PPO, GRPO and SDPO
  • Own training loops from rollout generation through to policy updates and checkpointing
  • Improve training throughput, GPU utilisation and memory efficiency
  • Profile and fix bottlenecks across distributed training
  • Investigate instability and differences between training and inference
  • Use real model failures to improve rewards, training data and overall performance

Skills

Programming
PyTorch
RL post-training
Distributed training
Experimental judgement
ML research & systems

Tools

vLLM
Ray
CUDA
Megatron-LM

Job description

Member of Technical Staff – Post-training / RL

Stealth AI Lab | London or Paris (Hybrid)

About

This role is about making models better after their initial training. You'll work on reinforcement learning and other post-training methods, while also improving the systems needed to run those experiments efficiently at scale.

The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.

You'll work across both the learning algorithms and the infrastructure underneath them. That means going from an RL experiment to a GPU profiler trace, finding what's limiting performance, and making sure systems improvements don't change the way the model learns.

What you\'ll do
  • Build and improve supervised fine-tuning, preference optimisation and RL methods
  • Work with approaches including PPO, GRPO and SDPO
  • Own training loops from rollout generation through to policy updates and checkpointing
  • Improve training throughput, GPU utilisation and memory efficiency
  • Profile and fix bottlenecks across distributed training
  • Investigate instability and differences between training and inference
  • Use real model failures to improve rewards, training data and overall performance
What you\'ll need
  • Strong programming and quantitative problem-solving skills
  • Hands-on experience with PyTorch and model training
  • Understanding of reinforcement learning or LLM post-training
  • Experience with distributed training and GPU systems
  • Strong experimental judgement
  • Ability to work comfortably across ML research and systems engineering
Optional
  • vLLM, SGLang, Ray, FSDP or Megatron-LM
  • CUDA, Triton, CuTE or GPU performance optimisation

Shortlisted candidates will be contacted within 48 hours.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff – RL Environments / Evals
Member of Technical Staff – RL Environments / Evals

Axiōma Search • Greater London

Hybrid
GBP 90,000 - 130,000
Member of Technical Staff (Post Training)
Member of Technical Staff (Post Training)

Inherentlabs • Greater London

On-site
GBP 80,000 - 100,000
Research Engineer, RL Scaling Science
Research Engineer, RL Scaling Science

Anthropic • Greater London

Hybrid
GBP 375,000 - 640,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Researcher, Training - London
Researcher, Training - London

OpenAI • Greater London

Hybrid
GBP 80,000 - 110,000
Relocation assistance
Hybrid work model
Member of Technical Staff - Research Scientist
Member of Technical Staff - Research Scientist

General Reasoning, Inc. • Greater London

On-site
GBP 70,000 - 90,000
Research Engineer / Scientist, Post-training - London
Research Engineer / Scientist, Post-training - London

H Company • Greater London

Hybrid
GBP 60,000 - 90,000
Competitive salary
Opportunities for professional growth
Collaborative and multicultural team environment
Research Engineer, Pretraining Scaling - London
Research Engineer, Pretraining Scaling - London

Anthropic • Greater London

On-site
GBP 250,000 - 435,000
Equity benefits
Visa sponsorship
Research Engineer, RL Scaling Science
Research Engineer, RL Scaling Science

Humanloop • Greater London

Hybrid
GBP 57,000 - 73,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+2
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Production AI Research Engineer
Production AI Research Engineer

Lovable • Greater London

On-site
GBP 120,000 - 180,000