Member of Technical Staff — Post-training / RL

Axiōma Search

London

On-site

NOK 1,154,000 - 1,538,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Axiōma Search is seeking a researcher to advance models after initial training, focusing on reinforcement learning and post-training methods. You will bridge ML research and the systems that run experiments at scale, from algorithmic improvements to GPU profiling and infrastructure tuning.

You'll collaborate across learning algorithms and the underlying systems, pushing throughput, reliability, and efficiency while keeping the learning signal intact.

Qualifications

  • Strong programming and quantitative problem-solving skills.
  • Hands-on experience with PyTorch and model training.
  • Understanding of reinforcement learning or LLM post-training.
  • Experience with distributed training and GPU systems.
  • Strong experimental judgement.
  • Ability to work across ML research and systems engineering.

Responsibilities

  • Build and improve supervised fine-tuning, preference optimisation and RL methods
  • Work with approaches including PPO, GRPO and SDPO
  • Own training loops from rollout generation through to policy updates and checkpointing
  • Improve training throughput, GPU utilisation and memory efficiency
  • Profile and fix bottlenecks across distributed training
  • Investigate instability and differences between training and inference
  • Use real model failures to improve rewards, training data and overall performance

Skills

Programming skills
PyTorch
Reinforcement learning
Distributed training
Experimentation
ML systems

Tools

Ray
Megatron-LM
vLLM

Job description

About

This role is about making models better after their initial training. You'll work on reinforcement learning and other post-training methods, while also improving the systems needed to run those experiments efficiently at scale.

The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.

You'll work across both the learning algorithms and the infrastructure underneath them. That means going from an RL experiment to a GPU profiler trace, finding what's limiting performance, and making sure systems improvements don't change the way the model learns.

What you'll do
  • Build and improve supervised fine-tuning, preference optimisation and RL methods
  • Work with approaches including PPO, GRPO and SDPO
  • Own training loops from rollout generation through to policy updates and checkpointing
  • Improve training throughput, GPU utilisation and memory efficiency
  • Profile and fix bottlenecks across distributed training
  • Investigate instability and differences between training and inference
  • Use real model failures to improve rewards, training data and overall performance
What you'll need
  • Strong programming and quantitative problem-solving skills
  • Hands-on experience with PyTorch and model training
  • Understanding of reinforcement learning or LLM post-training
  • Experience with distributed training and GPU systems
  • Strong experimental judgement
  • Ability to work comfortably across ML research and systems engineering
Optional
  • vLLM, SGLang, Ray, FSDP or Megatron-LM
  • CUDA, Triton, CuTE or GPU performance optimisation

Shortlisted candidates will be contacted within 48 hours.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer — RL & Post-Training
ML Systems Engineer — RL & Post-Training

Axiōma Search • London

On-site
NOK 1,154,000 - 1,538,000
AI Data Science Intern (UK)
AI Data Science Intern (UK)

TWG Global AI • London

On-site
NOK 589,000 - 710,000
Graduate Machine Learning Researcher
Graduate Machine Learning Researcher

Longshot Systems Ltd • London

On-site
GBP 50,000 - 70,000
Participation in uncapped bonus scheme
10% matched pension contributions
Private healthcare insurance
+3
ML engineer (LLM quantization & optimization)
ML engineer (LLM quantization & optimization)

ODS Serbia • Time

On-site
NOK 900,000 - 1,300,000
Investment Banking Intern (AI Evaluation)
Investment Banking Intern (AI Evaluation)

United States Digital Space LLC • London

On-site
NOK 225,000 - 393,000
Competitive salaries
London office in Liverpool Street
Referral bonus
Frontier Engineer
Frontier Engineer

Cognizant • Oslo

Hybrid
NOK 1,800,000 - 2,400,000
Python AI Engineer
Python AI Engineer

Square One Resources Limited • London

Hybrid
NOK 1,412,000 - 1,648,000
Member of Technical Staff, Applied AI
Member of Technical Staff, Applied AI

Latent Labs • London

On-site
NOK 1,557,430 - 2,465,932
Private health insurance
Pension contributions
Generous leave policies (including par
+1
AI developer
AI developer

Levato • Norway

On-site
NOK 1,000,000 - 1,400,000
Research Engineer (LLM Performance), London
Research Engineer (LLM Performance), London

Isomorphic Labs • London

On-site
NOK 877,000 - 1,504,000