ML Runtime Engineer: Scale Inference Engines

Fractile

West of England

Hybrid

GBP 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Private medical
Dental & Vision
Contributory Pension
25 days holiday + bank holidays
Life/Critical Illness Insurance

Job summary

Fractile is hiring a ML Runtime Engineer to integrate our AI accelerators with the latest inference frameworks and to build a scalable runtime stack. You will tackle challenges such as KV cache management and multi-user inference, focusing on transformer architectures in a collaborative environment.

The role requires deep experience in ML inference at scale, knowledge of paged attention and vLLM, and strong software engineering skills to maintain robust, high-performance systems.

Qualifications

  • Experience with ML inference at scale.
  • Deep understanding of paged attention and inference engines such as vLLM.
  • Familiarity with key components of the ML software ecosystem.
  • Strong software engineering skills and clean, maintainable systems.

Responsibilities

  • Integrate Fractile's AI accelerators with leading inference engines.
  • Research and implement KV cache management tailored to hardware.
  • Design and build a scalable reference inference engine.
  • Collaborate with the runtime team on transformer ML architectures.

Skills

ML inference at scale
Paged attention
vLLM familiarity
Software engineering

Education

Degree in Computer Science

Tools

Rust

Job description

Fractile is hiring a ML Runtime Engineer to integrate our AI accelerators with the latest inference frameworks and to build a scalable runtime stack. You will tackle challenges such as KV cache management and multi-user inference, focusing on transformer architectures in a collaborative environment.

The role requires deep experience in ML inference at scale, knowledge of paged attention and vLLM, and strong software engineering skills to maintain robust, high-performance systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
Mid-Level and Senior ML Runtime Engineer
Mid-Level and Senior ML Runtime Engineer

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
ML Runtime Engineer
ML Runtime Engineer

Fractile • West of England

Hybrid
GBP 90,000 - 150,000
Equity
Private medical
Dental & Vision
+3
Developer Experience Engineer - AI Inference Tools
Developer Experience Engineer - AI Inference Tools

Fractile • West of England

On-site
GBP 90,000 - 130,000
Equity ownership
Private medical
Contributory pension
+1
Senior ML Compiler Engineer for High-Speed AI Accelerators
Senior ML Compiler Engineer for High-Speed AI Accelerators

Fractile • Greater London

On-site
GBP 90,000 - 140,000
Equity
Private medical
Dental
+3
ML Inference & Serving Engineer
ML Inference & Serving Engineer

Google Inc. • Greater London

Hybrid
GBP 153,000 - 222,000
DevX Engineer: Build Tools to Run LLMs on AI Hardware
DevX Engineer: Build Tools to Run LLMs on AI Hardware

jobr.pro • Bristol

Hybrid
GBP 50,000 - 80,000
ML Infrastructure Engineer: Scalable LLM Serving Platform
ML Infrastructure Engineer: Scalable LLM Serving Platform

Neura Market • Greater London

On-site
GBP 110,000 - 165,000
Inference Systems Performance Engineer for AI Serving
Inference Systems Performance Engineer for AI Serving

adaption • Greater London

On-site
GBP 90,000 - 130,000
Flexible work
Lunch stipend
Well-Being benefits
Lead Inference Systems & Performance Engineer
Lead Inference Systems & Performance Engineer

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2