An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Axiōma Search is seeking a researcher to advance models after initial training, focusing on reinforcement learning and post-training methods. You will bridge ML research and the systems that run experiments at scale, from algorithmic improvements to GPU profiling and infrastructure tuning.
You'll collaborate across learning algorithms and the underlying systems, pushing throughput, reliability, and efficiency while keeping the learning signal intact.
This role is about making models better after their initial training. You'll work on reinforcement learning and other post-training methods, while also improving the systems needed to run those experiments efficiently at scale.
The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.
You'll work across both the learning algorithms and the infrastructure underneath them. That means going from an RL experiment to a GPU profiler trace, finding what's limiting performance, and making sure systems improvements don't change the way the model learns.
Shortlisted candidates will be contacted within 48 hours.