Inference Systems Performance Engineer for AI Serving

adaption

Greater London

On-site

GBP 90,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work
Lunch stipend
Well-Being benefits

Job summary

Adaption is seeking a senior ML systems engineer to own the cost and performance of our inference stack in a rapidly evolving environment. You will shape throughput, latency, and reliability by tuning caching, batching, and kernels while collaborating with the serving fleet engineers.

You will work with engines like vLLM, SGLang, and TensorRT-LLM, and build tools to measure where compute is spent. A strong background in Python and systems languages is essential, with a track record of real-world

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Inference infrastructure
Performance engineering
Python
C++
Rust
GPU performance

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
NCCL

Job description

Adaption is seeking a senior ML systems engineer to own the cost and performance of our inference stack in a rapidly evolving environment. You will shape throughput, latency, and reliability by tuning caching, batching, and kernels while collaborating with the serving fleet engineers.

You will work with engines like vLLM, SGLang, and TensorRT-LLM, and build tools to measure where compute is spent. A strong background in Python and systems languages is essential, with a track record of real-world

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer
Inference Performance Engineer

adaption • Greater London

On-site
GBP 90,000 - 130,000
Flexible work
Lunch stipend
Well-Being benefits
Lead Inference Systems & Performance Engineer
Lead Inference Systems & Performance Engineer

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2
Senior ML Systems Engineer — LLM Inference & Serving
Senior ML Systems Engineer — LLM Inference & Serving

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Senior DL Inference Engineer — GPU-Accelerated AI at Scale
Senior DL Inference Engineer — GPU-Accelerated AI at Scale

NVIDIA • United Kingdom

On-site
GBP 110,000 - 160,000
Competitive salaries
Extensive benefits package
Diversity & inclusion
Performance Engineer: AI Systems & Optimization
Performance Engineer: AI Systems & Optimization

CommonAI CIC • Cambridge

On-site
GBP 65,000 - 90,000
Collaborative environment
High impact in growing org
Competitive salary and pension
+3
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
Equity options
Competitive compensation
AI Inference Engineer | GPU-Scale Rust/Python | Equity
AI Inference Engineer | GPU-Scale Rust/Python | Equity

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
ML Inference & Serving Engineer
ML Inference & Serving Engineer

Google Inc. • Greater London

Hybrid
GBP 153,000 - 222,000
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000