Senior ML Engineer: GPU Inference & Low-Precision Training

Nebius

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive pay
Career growth
Flexibility and ownership
Collaborative culture
Impactful AI projects
International environment

Job summary

Nebius is building a high‑performance AI cloud platform enabling scalable inference and fine‑tuning across thousands of GPUs. You will contribute to optimizing throughput, latency and cost-per-token while advancing the state of the art in large language models.

The role focuses on research and production integrations, collaborating with global teams to push foundation models to hardware limits and drive measurable performance gains.

Qualifications

  • Proven expertise in ML theory and transformer models.
  • Strong Python coding skills and experience with PyTorch or TensorFlow.
  • Experience profiling GPU workloads (Nsight, PyTorch profiler, etc.).
  • Understanding of GPU memory hierarchy and compute/memory tradeoffs.
  • Familiar with MHA, RoPE, KV-cache, Flash Attention and quantisation.
  • Experience with modern DL frameworks and software engineering practices (CI/CD, tests).
  • Strong communication and leadership abilities.

Skills

Machine learning theory
Transformer architectures
Python
GPU profiling
Deep learning frameworks
CI/CD
Version control
Leadership
English communication

Tools

Nsight
PyTorch profiler
CUDA

Job description

Nebius is building a high‑performance AI cloud platform enabling scalable inference and fine‑tuning across thousands of GPUs. You will contribute to optimizing throughput, latency and cost-per-token while advancing the state of the art in large language models.

The role focuses on research and production integrations, collaborating with global teams to push foundation models to hardware limits and drive measurable performance gains.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Nebius • Greater London

On-site
GBP 100,000 - 140,000
Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
+3
Senior ML Infra Engineer: Scale GPU Clusters & Pipelines
Senior ML Infra Engineer: Scale GPU Clusters & Pipelines

Ellison Institute, LLC • Oxford

Hybrid
GBP 90,000 - 140,000
Competitive salary
25 days annual leave + 8 bank holidays
3 additional days between Christmas &.
+11
Senior DL Inference Engineer — GPU-Accelerated AI at Scale
Senior DL Inference Engineer — GPU-Accelerated AI at Scale

NVIDIA • United Kingdom

On-site
GBP 110,000 - 160,000
Competitive salaries
Extensive benefits package
Diversity & inclusion
Senior SRE - AI Inference Platform, Scale & Reliability
Senior SRE - AI Inference Platform, Scale & Reliability

Nebius • Greater London

On-site
GBP 100,000 - 140,000
Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
+3
Senior Serverless AI Engineer - GPUs & Cloud, Hybrid
Senior Serverless AI Engineer - GPUs & Cloud, Hybrid

Nebius • Greater London

Hybrid
GBP 120,000 - 170,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior ML Engineer — Production Inference & Optimization
Senior ML Engineer — Production Inference & Optimization

Hc1 • Greater London

Hybrid
GBP 90,000 - 150,000
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Staff ML Performance Engineer — Edge Inference Optimizer
Staff ML Performance Engineer — Edge Inference Optimizer

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
AI Compiler Engineer for Next-Gen GPU Inference
AI Compiler Engineer for Next-Gen GPU Inference

NVIDIA • Otley

On-site
GBP 90,000 - 150,000
Generous benefits package