Senior ML Engineer - Real-Time Inference at Scale

careers.bitkraft.vc - Jobboard

United Kingdom

On-site

GBP 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inworld is seeking a Full-Time ML Engineer to develop cutting-edge multimodal models. This role requires strong technical expertise in high-performance systems and experience with distributed systems, Kubernetes, and C++.

The base salary ranges from £140,000 to £200,000, plus equity and benefits. Candidates must have the legal right to work in the UK, as visa sponsorship is not available now.

Qualifications

  • Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
  • Hands-on experience with quantization, distillation, and caching strategies.
  • Proficiency in C++, CUDA, Rust, or Python for performance optimization.
  • Experience with Kubernetes, multi-GPU/multi-node inference.
  • Non-trivial systems programming projects or open-source contributions.
  • Ability to take a model from research to production.

Responsibilities

  • Develop and optimize multimodal models for real-time applications.
  • Design benchmarks and prototypes for unresolved problems.
  • Ensure performance, reliability, and latency as product features.

Skills

Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
Background in CS/Physics/Math

Education

PhD in Computer Science, Physics, Math, or equivalent

Tools

C++
CUDA
Rust
Python
Kubernetes
Ray

Job description

Inworld is seeking a Full-Time ML Engineer to develop cutting-edge multimodal models. This role requires strong technical expertise in high-performance systems and experience with distributed systems, Kubernetes, and C++.

The base salary ranges from £140,000 to £200,000, plus equity and benefits. Candidates must have the legal right to work in the UK, as visa sponsorship is not available now.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Realtime ML Systems Engineer - High-Performance Inference
Realtime ML Systems Engineer - High-Performance Inference

Inworld AI • United Kingdom

On-site
GBP 140,000 - 200,000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • United Kingdom

On-site
GBP 140,000 - 200,000
Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
Senior ML Performance Engineer - Real-Time Inference & Scale
Senior ML Performance Engineer - Real-Time Inference & Scale

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Senior C++ ML Inference Runtime Engineer
Senior C++ ML Inference Runtime Engineer

Semiconductor Engineering • United Kingdom

Remote
GBP 70,000 - 110,000
Senior ML Infra Architect for Large-Scale AI Simulations
Senior ML Infra Architect for Large-Scale AI Simulations

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Senior ML Engineer — Remote, Real-Time AI Scaling
Senior ML Engineer — Remote, Real-Time AI Scaling

High sky • Greater London

On-site
Fair salary based on interview results
Opportunity to enhance expertise
Comfortable schedule and fully remote work
Staff / Principal Research Scientist - UK
Staff / Principal Research Scientist - UK

BITKRAFT Ventures • United Kingdom

On-site
GBP 140,000 - 200,000
Senior ML Engineer — Real-Time Distributed Training
Senior ML Engineer — Real-Time Distributed Training

IMC • Greater London

On-site
GBP 90,000 - 140,000
Real-Time ML Systems Performance Engineer
Real-Time ML Systems Performance Engineer

Janestreet • Greater London

On-site
GBP 70,000 - 90,000