Senior ML Engineer: Production Inference & Optimization

Cloudflare

United Kingdom

Hybrid

GBP 90,000 - 160,000

Full time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
Commuter program
Healthcare
Pension plans
Maternity & paternity leave
Wellbeing programs

Job summary

Cloudflare seeks experienced ML Engineers to define and optimize how models run across its global network. You will work with systems engineers, product teams, and AI/ML engineers to bring models into production with low latency and high reliability.

You will build benchmarking frameworks, improve serving performance, and develop tooling that helps Cloudflare and its customers deploy AI at Internet scale. Strong Python and ML framework skills are essential.

Qualifications

  • Experience with large-scale inference serving frameworks or runtimes (SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp)
  • Hands-on experience with inference optimization techniques for GPUs/specialized accelerators
  • Experience building, optimizing, and operating ML models in production environments

Responsibilities

  • Define how ML models run across Cloudflare’s network.
  • Collaborate with systems engineers, product teams, and AI/ML engineers to productionize models with low latency
  • Benchmark models, improve serving performance, validate quality, and build tooling for AI at Internet scale
  • Develop, optimize, and productionize ML models for serverless inference platform focusing on performance and reliability
  • Build benchmarking/evaluation frameworks for latency, throughput, and cost across model families
  • Improve inference performance via quantization, batching, caching, compilation, and runtime tuning
  • Partner with teams to integrate models into Cloudflare’s distributed inference infrastructure
  • Drive model deployment workflows including validation, rollout safety, observability, and reliability
  • Mentor engineers and guide production ML engineering practices

Skills

Inference serving
Model optimization
Quantization
Batching
Caching
Python
ML frameworks

Tools

SGLang
vLLM
TensorRT-LLM
ONNX Runtime
Triton
llama.cpp

Job description

Cloudflare seeks experienced ML Engineers to define and optimize how models run across its global network. You will work with systems engineers, product teams, and AI/ML engineers to bring models into production with low latency and high reliability.

You will build benchmarking frameworks, improve serving performance, and develop tooling that helps Cloudflare and its customers deploy AI at Internet scale. Strong Python and ML framework skills are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer: Production Inference & Optimization
Senior ML Engineer: Production Inference & Optimization

Cloudflare • Greater London

Hybrid
GBP 110,000 - 140,000
Senior ML Engineer: Inference & Latency Optimization
Senior ML Engineer: Inference & Latency Optimization

Nebius Group • Greater London

On-site
GBP 90,000 - 140,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior ML Engineer: Production-Scale AI & Cloud Leader
Senior ML Engineer: Production-Scale AI & Cloud Leader

Faculty Science Limited • Greater London

Hybrid
GBP 90,000 - 130,000
Unlimited Annual Leave Policy
Private healthcare and dental
Enhanced parental leave
+3
Senior ML Engineer — Production-Grade AI for Finance
Senior ML Engineer — Production-Grade AI for Finance

NLP PEOPLE • Milton Keynes

Hybrid
GBP 75,000 - 110,000
Unlimited Annual Leave
Private healthcare
Enhanced parental leave
+3
Senior ML Engineer – Production Systems & Impact
Senior ML Engineer – Production Systems & Impact

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 150,000
Unlimited Annual Leave Policy
Private healthcare and dental
Enhanced parental leave
+3
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Cloudflare • Greater London

On-site
GBP 110,000 - 140,000
ML Infrastructure Engineer - Scale AI Training & Inference
ML Infrastructure Engineer - Scale AI Training & Inference

Isomorphic Labs • City of Westminster

On-site
GBP 90,000 - 130,000
Production ML Engineer - Scale Models & LLM Fine-Tunes
Production ML Engineer - Scale Models & LLM Fine-Tunes

AutoThink Group • Rochdale

Hybrid
GBP 90,000 - 130,000
Senior Data Scientist — AI/ML Platform & GenAI Innovator
Senior Data Scientist — AI/ML Platform & GenAI Innovator

Cloudflare • United Kingdom

On-site
GBP 80,000 - 120,000
Competitive pay
Generous vacation policy
Maternity & paternity leave
+3
Lead ML Ops Engineer — Production AI Platform (Hybrid)
Lead ML Ops Engineer — Production AI Platform (Hybrid)

Artificial Intelligence Jobs • Greater London

Hybrid
GBP 75,000 - 85,000
Annual bonus (10%)