Senior ML Engineer — AI Inference & Performance

Nebius Group

London (KY)

On-site

USD 160,000 - 260,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Career growth
Flexibility
Collaborative culture
Impactful AI projects
International teams

Job summary

Nebius is building a fast, cost-efficient AI cloud platform. We seek a Senior Machine Learning Engineer to own model and endpoint optimization from artefacts to production deployment.

You will work across model internals, inference engines, and serving architectures, aiming to improve latency, throughput, memory efficiency, GPU utilization and cost per token while preserving model quality. This hands-on role involves benchmarking, reproducible results, and close collaboration with kernel and

Qualifications

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM/VLM or high-throughput transformer inference systems.
  • Practical knowledge of inference stacks such as vLLM, SGLang, TensorRT-LLM, Triton, NVIDIA Dynamo, Ray Serve, or KServe.
  • Understanding of transformer bottlenecks (KV cache, attention, memory bandwidth) and latency tradeoffs.
  • Ability to reason quantitatively about latency, throughput, quality, and cost.

Responsibilities

  • Own optimization work for specific model families, customer endpoints, or serving backends.
  • Run engine comparisons and recommend practical serving configurations for workloads.
  • Debug model quality or performance regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput and cost per token.
  • Deploy, configure, benchmark, and extend inference engines (e.g., vLLM, TensorRT-LLM, Triton).
  • Build and productionize model compression workflows (quantization, distillation, low-bit serving).
  • Implement speculative decoding, KV-cache optimization, and continuous batching.
  • Create reproducible benchmarks for latency, tokens per second per GPU, and memory.
  • Collaborate with kernel and platform engineers to diagnose bottlenecks.
  • Write design docs, performance reports, and rollout plans.

Skills

Python
PyTorch
LLM inference
VLM inference
Transformer optimization

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
KServe
Ray Serve

Job description

Nebius is building a fast, cost-efficient AI cloud platform. We seek a Senior Machine Learning Engineer to own model and endpoint optimization from artefacts to production deployment.

You will work across model internals, inference engines, and serving architectures, aiming to improve latency, throughput, memory efficiency, GPU utilization and cost per token while preserving model quality. This hands-on role involves benchmarking, reproducible results, and close collaboration with kernel and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
GPU Systems Engineer for AI Inference & Performance
GPU Systems Engineer for AI Inference & Performance

Nebius • United States

Remote
USD 180,000 - 260,000
Competitive compensation
Career growth
Flexible ownership & autonomy
+3
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI Research Scientist — Architecting Efficient Models
Senior AI Research Scientist — Architecting Efficient Models

Nebius • United States

Remote
USD 150,000 - 280,000
Competitive compensation
Career growth
Flexibility and ownership
+3
GPU ML Benchmarking Engineer for Next-Gen AI Infra
GPU ML Benchmarking Engineer for Next-Gen AI Infra

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior ML Inference Engineer – High-Performance Serving
Senior ML Inference Engineer – High-Performance Serving

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior AI/ML Inference Engineer (Neuron)
Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
Senior ML Scientist - Inference & Hardware Acceleration
Senior ML Scientist - Inference & Hardware Acceleration

Netskope • Santa Clara (CA)

On-site
USD 182,500 - 260,500
Senior ML Solutions Architect – Remote Token Platform (LLM)
Senior ML Solutions Architect – Remote Token Platform (LLM)

Nebius • United States

On-site
USD 210,000 - 260,000
Health Insurance
401(k) Plan
Parental Leave
+2