Lead Inference Systems & Performance Engineer

United States Digital Space LLC

Greater London

On-site

GBP 101,000 - 192,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity & Ownership
Private healthcare
Visa sponsorship
Relocation benefits
London office

Job summary

Intelligent Systems Company in London seeks a senior Inference Platform Engineer to own end‑to‑end performance for our heterogeneous inference stack. You will drive KV cache strategies, batching, memory management, and multi‑node scheduling to scale models and accelerate silicon efficiency.

The role requires deep knowledge of LLM internals, strong systems engineering, and hands‑on debugging across GPU and Linux. We offer visa sponsorship and in‑person work at our London office.

Qualifications

  • Deep understanding of LLM inference internals: KV cache lifecycle, memory management, attention mechanisms, and serving architectures.
  • Strong systems engineering background with proven experience optimising distributed GPU workloads.
  • Proficiency in C++, CUDA, Python, Rust, or similar - and the instinct to go low‑level when it matters.
  • Hands‑on debugging skills across GPU, networking, and Linux systems - able to work from first principles with limited tooling.
  • Experience building or significantly optimising production‑grade, high‑throughput model serving stacks.
  • Multi‑GPU and multi‑node inference optimisation using NCCL, NVLink, RDMA, InfiniBand, or RoCE.
  • GPU memory profiling, CUDA or Triton kernel optimisation.
  • Linux performance analysis and optimisation

Responsibilities

  • Design and optimise inference serving systems across heterogeneous multi‑GPU and multi‑node environments.
  • Own KV cache lifecycle management, batching strategies, and memory allocation to maximise throughput and minimise latency.
  • Profile and tune GPU kernels, identify bottlenecks across compute, memory, and network, and implement targeted optimisations.
  • Build and improve scheduling logic for continuous batching, disaggregated prefill/decode, and speculative decoding.
  • Work with networking primitives - NCCL, NVLink, RDMA, InfiniBand, RoCE - to optimise communication across distributed inference workloads.
  • Develop tooling for performance visibility, regression detection, and benchmarking across hardware configurations.

Skills

LLM inference internals
KV cache lifecycle
Memory management
High-performance systems
Distributed GPU workloads
C++
CUDA
Python
Rust

Tools

C++
CUDA
Python
Rust

Job description

Intelligent Systems Company in London seeks a senior Inference Platform Engineer to own end‑to‑end performance for our heterogeneous inference stack. You will drive KV cache strategies, batching, memory management, and multi‑node scheduling to scale models and accelerate silicon efficiency.

The role requires deep knowledge of LLM internals, strong systems engineering, and hands‑on debugging across GPU and Linux. We offer visa sponsorship and in‑person work at our London office.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Systems Engineer – Multi-Node GPU
Senior Inference Systems Engineer – Multi-Node GPU

Callosum • Greater London

On-site
GBP 120,000 - 160,000
Equity
Private healthcare
Visa sponsorship
+2
Inference Systems Performance Engineer for AI Serving
Inference Systems Performance Engineer for AI Serving

adaption • Greater London

On-site
GBP 90,000 - 130,000
Flexible work
Lunch stipend
Well-Being benefits
LLM Inference Performance Engineer
LLM Inference Performance Engineer

G-Research • Greater London

On-site
GBP 90,000 - 150,000
Discretionary bonus
35 days leave
Pension contributions
+2
Inference System & Performance - Member of Technical Staff
Inference System & Performance - Member of Technical Staff

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2
Staff Inference Platform Engineer — Low-Latency GPU, Kubernetes
Staff Inference Platform Engineer — Low-Latency GPU, Kubernetes

CoreWeave Europe • Greater London

On-site
GBP 120,000 - 190,000
Medical Insurance
Pension Plan
Life Insurance
+1
Senior AI Infra & DS Engineer — London
Senior AI Infra & DS Engineer — London

LinuxRecruit • Greater London

On-site
GBP 100,000 - 140,000
Inference Performance Engineer
Inference Performance Engineer

adaption • Greater London

On-site
GBP 90,000 - 130,000
Flexible work
Lunch stipend
Well-Being benefits
Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
Hardware-Aware Inference Engine Engineer
Hardware-Aware Inference Engine Engineer

Uncover • Greater London

Hybrid
GBP 110,000 - 170,000
Competitive salary
Equity & Ownership
Private healthcare
+1
Staff Software Engineer, Inference — Scalable AI Cloud
Staff Software Engineer, Inference — Scalable AI Cloud

United States Digital Space LLC • Greater London

On-site
GBP 120,000 - 180,000
Family-level Medical Insurance
Dental Insurance
Generous Pension Contribution
+1