Research Engineer (Inference & Serving)

Axiōma Search

Greater London

On-site

GBP 90,000 - 120,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Axiōma Search is building a state-of-the-art multimodal agent inference stack in production, spanning from engine layers to serving architecture. You will help design and operate systems for low latency, high throughput, and cost efficiency while collaborating with research and engineering teams.

The role focuses on research-driven production ML and scalable infrastructure, with exposure to modern transformers and multimodal architectures in a fast-paced, VC-backed setting.

Qualifications

  • Strong software engineering fundamentals with production experience.
  • Proficient in Python and at least one systems language (Rust, C++, or Go).
  • Hands-on experience with PyTorch or JAX in an industry setting.
  • Experience with inference frameworks: vLLM, SGLang, TensorRT-LLM.
  • Solid distributed systems and production ML infrastructure experience.
  • Knowledge of modern ML including transformers and multimodal architectures.

Responsibilities

  • Build and operate the inference stack serving multimodal agentic models in production.
  • Improve latency, throughput, and cost across the serving stack.
  • Research and implement inference techniques tailored to agent workloads.
  • Co-design with the models team on training-time decisions that affect inference behaviour.
  • Evaluate inference frameworks and hardware platforms and feed findings back into roadmap decisions.
  • Stay current with advances in inference, model serving, and accelerator technology.

Skills

Python
Rust
C++
Go
PyTorch
JAX
Distributed systems
Multimodal ML

Tools

vLLM
SGLang
TensorRT-LLM
CUDA

Job description

About

Serving a multimodal agent model in production is a different problem to serving a standard LLM. Context length, tool calls, and computer-use workloads create constraints that require co-designing the inference stack with the model team - not just bolting on a serving framework after the fact.

This is a VC-backed challenger lab building state-of-the-art computer-use agents. The inference team owns the full stack from engine layer (vLLM, SGLang) through to serving architecture (disaggregated inference, intelligent routing).

The team operates at the intersection of research and production - translating cutting-edge techniques directly into the systems behind live agent products.

What you'll do
  • Build and operate the inference stack serving multimodal agentic models in production
  • Improve latency, throughput, and cost across the serving stack
  • Research and implement inference techniques tailored to agent workloads
  • Co-design with the models team on training-time decisions that affect inference behaviour
  • Evaluate inference frameworks and hardware platforms and feed findings back into roadmap decisions
  • Stay current with advances in inference, model serving, and accelerator technology
What you'll need
  • Strong software engineering fundamentals and a solid production track record
  • Proficient in Python and at least one systems language - Rust, C++, or Go
  • Hands-on experience with PyTorch or JAX in an industry setting
  • Experience with inference frameworks: vLLM, SGLang, TensorRT-LLM
  • Solid distributed systems fundamentals and experience operating production ML infrastructure
  • Working knowledge of modern ML including transformers and multimodal architectures
Optional Bonus
  • Research engagement: advanced degree with research output, top-tier publications (NeurIPS, ICML, MLSys, OSDI), or open-source contributions
  • GPU kernel work - CUDA, Triton, or similar
  • Experience with quantisation, speculative decoding, disaggregated inference, or KV-cache compression
  • Shortlisted candidates will be contacted within 48 hours.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity
Senior Software Engineer (Inference)
Senior Software Engineer (Inference)

AssemblyAI • Greater London

Remote
GBP 90,000 - 150,000
Home office stipend
Equity grant
Premium medical, dental, vision plans
+4
Multimodal Inference & Serving Engineer
Multimodal Inference & Serving Engineer

Axiōma Search • Greater London

On-site
GBP 90,000 - 120,000
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Research Engineer (Post-Training)
Research Engineer (Post-Training)

Axiōma Search • Greater London

Hybrid
GBP 90,000 - 140,000
AI Infrastructure Engineer, Serving Platform
AI Infrastructure Engineer, Serving Platform

scaleai • Greater London

On-site
GBP 110,000 - 160,000
Research Engineer (Agentic Models)
Research Engineer (Agentic Models)

JetBrains • Greater London

On-site
GBP 60,000 - 80,000
Research Engineer / Scientist, Post-training - London
Research Engineer / Scientist, Post-training - London

H Company • Greater London

Hybrid
GBP 60,000 - 90,000
Competitive salary
Opportunities for professional growth
Collaborative and multicultural team environment
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • United Kingdom

On-site
GBP 140,000 - 200,000
AI Infrastructure Engineer, Serving Platform
AI Infrastructure Engineer, Serving Platform

Mat Vin • Greater London

Hybrid
GBP 90,000 - 150,000