Senior Inference Systems Engineer – Multi-Node GPU

Callosum

Greater London

On-site

GBP 120,000 - 160,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Private healthcare
Visa sponsorship
Relocation benefits
London office

Job summary

Callosum, based in London, is seeking an experienced infrastructure-focused engineer to own end-to-end performance for its heterogeneous inference platforms. The role spans KV cache strategies, batching, memory management and multi-node scheduling to drive speed and efficiency across a growing model portfolio.

You will design, profile, and optimise GPU kernels and networking paths, build tooling for visibility, and collaborate across the stack to scale hardware and models.

Qualifications

  • Experience with KV cache lifecycle, memory management, attention mechanisms, and serving architectures.
  • Proven track record optimising distributed GPU workloads.
  • Proficiency in C++, CUDA, Python, Rust, or similar.
  • Hands-on debugging across GPU, networking, and Linux systems.

Responsibilities

  • Design and optimise inference serving systems across heterogeneous multi-GPU and multi-node environments.
  • Own KV cache lifecycle management, batching strategies, and memory allocation to maximise throughput and minimise latency.
  • Profile and tune GPU kernels, identify bottlenecks across compute, memory, and network, and implement targeted optimisations.
  • Build and improve scheduling logic for continuous batching, disaggregated prefill/decode, and speculative decoding.
  • Work with networking primitives - NCCL, NVLink, RDMA, InfiniBand, RoCE - to optimise communication across distributed inference workloads.
  • Develop tooling for performance visibility, regression detection, and benchmarking across hardware configurations.

Skills

LLM inference
Systems engineering
C++
CUDA
Python
Rust
Linux

Tools

NCCL
NVLink
RDMA
InfiniBand
RoCE

Job description

Callosum, based in London, is seeking an experienced infrastructure-focused engineer to own end-to-end performance for its heterogeneous inference platforms. The role spans KV cache strategies, batching, memory management and multi-node scheduling to drive speed and efficiency across a growing model portfolio.

You will design, profile, and optimise GPU kernels and networking paths, build tooling for visibility, and collaborate across the stack to scale hardware and models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Inference Systems & Performance Engineer
Lead Inference Systems & Performance Engineer

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2
Staff Engineer, AI Inference Performance & Deployment
Staff Engineer, AI Inference Performance & Deployment

Uncover • Greater London

Hybrid
GBP 110,000 - 160,000
Competitive salary
Equity & ownership
Private healthcare
+3
Inference System & Performance - Member of Technical Staff
Inference System & Performance - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 160,000
Equity
Private healthcare
Visa sponsorship
+2
Hardware-Aware Inference Engine Engineer
Hardware-Aware Inference Engine Engineer

Uncover • Greater London

Hybrid
GBP 110,000 - 170,000
Competitive salary
Equity & Ownership
Private healthcare
+1
Staff Systems Software Engineer - Heterogeneous AI Accel
Staff Systems Software Engineer - Heterogeneous AI Accel

Uncover • Greater London

Hybrid
GBP 120,000 - 180,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation benefits
+1
Senior GPU HPC Engineer: InfiniBand & KVM Optimization
Senior GPU HPC Engineer: InfiniBand & KVM Optimization

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Staff Inference Platform Engineer — Low-Latency GPU, Kubernetes
Staff Inference Platform Engineer — Low-Latency GPU, Kubernetes

CoreWeave Europe • Greater London

On-site
GBP 120,000 - 190,000
Medical Insurance
Pension Plan
Life Insurance
+1
Inference Performance & Deployment - Member of Technical Staff
Inference Performance & Deployment - Member of Technical Staff

Uncover • Greater London

Hybrid
GBP 110,000 - 160,000
Competitive salary
Equity & ownership
Private healthcare
+3
Inference System & Performance - Member of Technical Staff
Inference System & Performance - Member of Technical Staff

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2
Senior GPU Systems Engineer - Clusters & Platform Infra
Senior GPU Systems Engineer - Clusters & Platform Infra

Radley James • Greater London

On-site
GBP 20,000 - 40,000