Hardware-Aware Inference Engine Engineer

Uncover

Greater London

Hybrid

GBP 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity & Ownership
Private healthcare
Visa sponsorship and relocation

Job summary

Callosum is seeking a seasoned ML inference engineer to push heterogeneous hardware orchestration beyond GPUs. You will work on SGLang and vLLM, extending them to run efficiently across diverse accelerators, focusing on scheduling, memory and execution.

Based in London, you will collaborate with an Accelerator Systems Software engineer, contribute upstream, and maintain internal forks while scaling production inference on mixed hardware.

Qualifications

  • Deep familiarity with SGLang and vLLM, or comparable inference serving frameworks, including scheduler design, memory management and execution pipelines.
  • Strong background in high-performance Python and C++/CUDA systems, particularly in ML inference.
  • Experience designing or implementing parallelism strategies for large model serving.
  • Understanding of disaggregated serving architectures and tradeoffs in separating modules of a workflow.
  • Proven ability to contribute to fast-moving open source codebases with evolving APIs and design conventions.

Responsibilities

  • Contribute upstream to SGLang and vLLM, and maintain internal forks where our requirements diverge.
  • Improve hardware-awareness within inference engines so that scheduling, memory management, and execution adapt to the capabilities of the underlying accelerator.
  • Design and implement bespoke parallelism and disaggregation strategies that go beyond default configurations to better exploit heterogeneous hardware.
  • Work closely with an Accelerator Systems Software engineer to ensure engine-level abstractions map cleanly onto diverse hardware capabilities.

Skills

SGLang familiarity / vLLM
Scheduler design
Memory management
C++ / CUDA
Large model serving

Tools

CUDA
Linux

Job description

Callosum is seeking a seasoned ML inference engineer to push heterogeneous hardware orchestration beyond GPUs. You will work on SGLang and vLLM, extending them to run efficiently across diverse accelerators, focusing on scheduling, memory and execution.

Based in London, you will collaborate with an Accelerator Systems Software engineer, contribute upstream, and maintain internal forks while scaling production inference on mixed hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engine Development - Member of Technical Staff
Inference Engine Development - Member of Technical Staff

Uncover • Greater London

Hybrid
GBP 110,000 - 170,000
Competitive salary
Equity & Ownership
Private healthcare
+1
Staff Systems Software Engineer - Heterogeneous AI Accel
Staff Systems Software Engineer - Heterogeneous AI Accel

Uncover • Greater London

Hybrid
GBP 120,000 - 180,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation benefits
+1
Staff Engineer, AI Inference Performance & Deployment
Staff Engineer, AI Inference Performance & Deployment

Uncover • Greater London

Hybrid
GBP 110,000 - 160,000
Competitive salary
Equity & ownership
Private healthcare
+3
Senior Inference Systems Engineer – Multi-Node GPU
Senior Inference Systems Engineer – Multi-Node GPU

Callosum • Greater London

On-site
GBP 120,000 - 160,000
Equity
Private healthcare
Visa sponsorship
+2
Accelerator Systems Software - Member of Technical Staff
Accelerator Systems Software - Member of Technical Staff

Uncover • Greater London

Hybrid
GBP 120,000 - 180,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation benefits
+1
Lead Heterogeneous Interconnects & RDMA Networking
Lead Heterogeneous Interconnects & RDMA Networking

Callosum • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity & Ownership
Private healthcare
+3
Inference System & Performance - Member of Technical Staff
Inference System & Performance - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 160,000
Equity
Private healthcare
Visa sponsorship
+2
Lead Inference Systems & Performance Engineer
Lead Inference Systems & Performance Engineer

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2
Networking & Interconnect Systems - Member of Technical Staff
Networking & Interconnect Systems - Member of Technical Staff

Callosum • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity & Ownership
Private healthcare
+3
Inference Performance & Deployment - Member of Technical Staff
Inference Performance & Deployment - Member of Technical Staff

Uncover • Greater London

Hybrid
GBP 110,000 - 160,000
Competitive salary
Equity & ownership
Private healthcare
+3