Mid-Level and Senior ML Runtime Engineer

Fractile

Bristol

Hybrid

GBP 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Hybrid working
Visible and valued contributions

Job summary

Fractile is looking for a Senior ML Runtime Engineer in the UK to integrate AI acceleration hardware with inference frameworks. Join a small expert team to tackle complex problems like KV cache management and scalable multi-user inference.

This hybrid role offers competitive salary and equity, along with a culture that values learning and collaboration. You'll have the opportunity to shape the runtime stack of cutting-edge AI technologies while working in offices located in London and Bristol.

Qualifications

  • Solid experience with ML inference at scale, including multi-user serving.
  • A deep understanding of paged attention and inference engines such as vLLM.
  • Strong software engineering skills and an instinct for clean, maintainable systems.

Responsibilities

  • Integrate Fractile's AI acceleration hardware with leading inference engines.
  • Research KV cache management technologies and build proof-of-concept implementations.
  • Work closely with the runtime team to design and build a scalable inference engine.
  • Focus primarily on the transformer ML architecture.

Skills

ML inference at scale
Paged attention
Inference engines like vLLM
Software engineering

Education

Degree in Computer Science or related field

Tools

Rust

Job description

About Fractile

We’re taking a revolutionary approach to computing — building AI acceleration hardware that runs the world’s largest language models 100× faster than existing systems. Our team works at the cutting edge of both hardware and software AI development, and we’re growing fast.

Fractile

  • London or Bristol
  • Full-time
  • Hybrid
The Role

We’re looking for a Senior ML Runtime Engineer to help us integrate Fractile’s AI accelerators with the latest inference frameworks and build the runtime stack that makes them fly. You’ll work on genuinely hard problems — KV cache management, scalable multi‑user inference, and the internals of transformer model execution — alongside a collaborative team that values curiosity and rigor equally.

This is a hybrid role, with offices in London and Bristol — your choice of base.

What You’ll Do
  • Integrate Fractile’s AI acceleration hardware with leading inference engines including vLLM and SGLang
  • Research KV cache management technologies (including paged attention) and build proof‑of‑concept implementations tailored to our hardware
  • Work closely with the runtime team to design and build a scalable, bare‑bones reference inference engine
  • Focus primarily on the transformer ML architecture
  • Share your expertise to help shape the direction of our runtime stack
What We’re Looking For

We care most about depth of knowledge and a genuine interest in the problem space. You’ll be a strong fit if you have:

  • Solid experience with ML inference at scale, including multi‑user serving
  • A deep understanding of paged attention and inference engines such as vLLM
  • Familiarity with key components of the ML software ecosystem
  • Strong software engineering skills and an instinct for clean, maintainable systems
Bonus Points
  • Experience with Rust
  • Having built your own inference engine from scratch
  • A degree in Computer Science or a related field
Why Fractile
  • Work on one of the most technically ambitious projects in AI infrastructure
  • A small, expert team where your contributions are visible and valued
  • Hybrid working — split your time between home and our London or Bristol office
  • Competitive salary and equity
  • A culture that values learning, directness, and collaboration

Fractile is committed to building a diverse and inclusive team. We welcome applications from people of all backgrounds and actively encourage candidates from underrepresented groups to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Runtime Engineer
ML Runtime Engineer

Fractile • West of England

Hybrid
GBP 90,000 - 150,000
Equity
Private medical
Dental & Vision
+3
Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
Developer Experience Engineer New London
Developer Experience Engineer New London

Fractile Ltd • Greater London

Hybrid
GBP 70,000 - 120,000
Equity
Hybrid work model
Competitive salary
Senior ML Compiler Engineer
Senior ML Compiler Engineer

Fractile • Greater London

On-site
GBP 90,000 - 140,000
Equity
Private medical
Dental
+3
Rust Software Engineer
Rust Software Engineer

Fractile • West of England

On-site
GBP 90,000 - 140,000
Competitive salary
Equity
Private Medical
+3
Developer Experience Engineer
Developer Experience Engineer

Fractile • Greater London

Hybrid
GBP 70,000 - 90,000
Rust Software Engineer
Rust Software Engineer

Fractile • Greater London

On-site
GBP 70,000 - 120,000
Equity & Ownership
Private Medical
Dental and Vision
+3
Modelling Engineer
Modelling Engineer

Fractile • Greater London

On-site
GBP 70,000 - 110,000
Competitive salary
Equity & Ownership
Benefits: Private Medical, Dental and…
+1
Senior Systems Software Engineer
Senior Systems Software Engineer

Fractile • England

On-site
GBP 60,000 - 80,000
Senior Linux Kernel Driver Engineer
Senior Linux Kernel Driver Engineer

Fractile • Bristol

Hybrid
GBP 60,000 - 80,000