Senior ML Engineer: Inference & Latency Optimization

Nebius Group

Greater London

On-site

GBP 90,000 - 140,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Career growth and learning
Flexibility and ownership
Collaborative culture
Impactful AI projects
International environment

Job summary

Nebius is seeking a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts to production deployment on our Applied AI team. You will improve latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality.

This hands-on role involves diagnosing complex serving problems, evaluating configurations, and delivering measurable production improvements in collaboration with kernel and platform engineers.

Qualifications

  • Strong Python and PyTorch engineering skills with production exposure.
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
  • Practical knowledge of modern inference stacks including vLLM, SGLang, TensorRT-LLM, Triton, or similar.
  • Understanding of transformer bottlenecks: KV cache, attention, memory bandwidth, batching, and long-context serving.
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs.
  • Strong communication and collaboration across diverse teams.

Responsibilities

  • Own optimization work for specific model families, customer endpoints, or serving backends.
  • Run engine comparisons and recommend practical serving configurations for workloads.
  • Debug model quality or performance regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton, Dynamo or similar systems.
  • Build production-ready model-compression workflows including quantization and distillation.
  • Implement or integrate speculative decoding, KV-cache optimization, prefix caching, and continuous batching.
  • Create reproducible benchmark harnesses for latency, tokens per second per GPU, and cost per token.
  • Collaborate with kernel and platform engineers to diagnose bottlenecks.

Skills

Python
PyTorch
LLM inference
Performance optimization
Systems thinking
Communication

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
Ray Serve
KServe

Job description

Nebius is seeking a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts to production deployment on our Applied AI team. You will improve latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality.

This hands-on role involves diagnosing complex serving problems, evaluating configurations, and delivering measurable production improvements in collaboration with kernel and platform engineers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer: GPU Inference & Low-Precision Training
Senior ML Engineer: GPU Inference & Low-Precision Training

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior ML Engineer: Production Inference & Optimization
Senior ML Engineer: Production Inference & Optimization

Cloudflare • United Kingdom

Hybrid
GBP 90,000 - 160,000
Stock options
Commuter program
Healthcare
+3
Senior ML Engineer: Production Inference & Optimization
Senior ML Engineer: Production Inference & Optimization

Cloudflare • Greater London

Hybrid
GBP 110,000 - 140,000
Senior ML Performance Engineer - Real-Time Inference & Scale
Senior ML Performance Engineer - Real-Time Inference & Scale

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Staff ML Performance Engineer — Edge Inference Optimizer
Staff ML Performance Engineer — Edge Inference Optimizer

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
Senior Backend Engineer – Low-Latency AI Search
Senior Backend Engineer – Low-Latency AI Search

Nebius • Greater London

On-site
GBP 110,000 - 160,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Nebius Group • Greater London

On-site
GBP 90,000 - 140,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Real-Time ML Inference Engineer
Senior Real-Time ML Inference Engineer

BITKRAFT Ventures • United Kingdom

On-site
GBP 140,000 - 200,000
Senior ML Engineer — Optimizing AI Models at Scale
Senior ML Engineer — Optimizing AI Models at Scale

EngineersOfAI • Greater London

On-site
GBP 110,000 - 140,000
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000