Realtime ML Systems Engineer - High-Performance Inference

Inworld AI

United Kingdom

On-site

GBP 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI research lab is seeking talented individuals to develop sophisticated multimodal models and optimization techniques. The ideal candidate will have a PhD or equivalent experience in CS, Physics, or Math, with proficiency in high-performance systems and distributed scaling solutions. Responsibilities include taking models into production and ensuring performance and reliability across thousands of queries per second. The position offers a base salary of £140,000 – £200,000, along with equity and benefits.

Qualifications

  • Deep understanding of modern serving frameworks and optimization techniques.
  • Hands-on experience with model quantization and caching strategies.
  • Proficiency in performance optimization on NVIDIA GPUs.

Responsibilities

  • Work on agentic systems and multimodal inference at scale.
  • Take models from the research team and optimize their serving.
  • Ensure reliability in production environments.

Skills

Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
Background in CS, Physics, Math

Education

PhD in CS, Physics, Math or equivalent experience

Tools

C++
CUDA
Rust
Python
Kubernetes
Ray

Job description

A leading AI research lab is seeking talented individuals to develop sophisticated multimodal models and optimization techniques. The ideal candidate will have a PhD or equivalent experience in CS, Physics, or Math, with proficiency in high-performance systems and distributed scaling solutions. Responsibilities include taking models into production and ensuring performance and reliability across thousands of queries per second. The position offers a base salary of £140,000 – £200,000, along with equity and benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - Real-Time Inference at Scale
Senior ML Engineer - Real-Time Inference at Scale

careers.bitkraft.vc - Jobboard • United Kingdom

On-site
GBP 140,000 - 200,000
Senior ML Performance Engineer - Real-Time Inference & Scale
Senior ML Performance Engineer - Real-Time Inference & Scale

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Senior ML Infra Architect for Large-Scale AI Simulations
Senior ML Infra Architect for Large-Scale AI Simulations

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Research Engineer: Scalable Production ML Systems
Research Engineer: Scalable Production ML Systems

Anthropic • Greater London

On-site
GBP 260,000 - 630,000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • United Kingdom

On-site
GBP 140,000 - 200,000
Lead ML Infra Engineer for Scalable Physics AI
Lead ML Infra Engineer for Scalable Physics AI

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Staff Software Engineer, Large-Scale Inference
Staff Software Engineer, Large-Scale Inference

Anthropic • Greater London

Hybrid
GBP 80,000 - 100,000
Competitive compensation
Generous vacation
Flexible working hours
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
Inference Systems Performance Engineer for AI Serving
Inference Systems Performance Engineer for AI Serving

adaption • Greater London

On-site
GBP 90,000 - 130,000
Flexible work
Lunch stipend
Well-Being benefits