Senior ML Engineer: Production Inference & Optimization

Cloudflare

Greater London

Hybrid

GBP 110,000 - 140,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cloudflare in London is seeking an ML engineer to define how machine learning models run across Cloudflare’s global network. You’ll collaborate with systems engineers, product teams, hardware partners, and AI/ML engineers to deploy models with low latency, strong reliability, and efficient resource use in production.

You’ll focus on benchmarking, evaluation, and tooling to help Cloudflare and its customers ship AI applications at internet scale.

Qualifications

  • Experience building, optimizing, and operating ML models in production.
  • Proficiency with Python and ML frameworks (PyTorch, TensorFlow, JAX).
  • Hands-on inference optimization for large-scale models (quantization, batching, caching).
  • Experience with model serving runtimes (TensorRT-LLM, ONNX Runtime, Triton).
  • Familiarity with LLMs, speech/vision models, embeddings, multimodal modeling.

Responsibilities

  • Develop, optimize, and productionize ML models for Cloudflare’s serverless inference platform.
  • Build benchmarking/evaluation frameworks for latency, throughput, and cost efficiency across model families.
  • Improve inference performance via quantization, batching, caching, and runtime tuning.

Skills

Python
PyTorch
TensorFlow
JAX
LLMs
Inference optimization
Model deployment
Performance tuning
Mentoring
Distributed systems

Tools

SGLang
vLLM
TensorRT-LLM
ONNX Runtime
Triton
llama.cpp

Job description

Cloudflare in London is seeking an ML engineer to define how machine learning models run across Cloudflare’s global network. You’ll collaborate with systems engineers, product teams, hardware partners, and AI/ML engineers to deploy models with low latency, strong reliability, and efficient resource use in production.

You’ll focus on benchmarking, evaluation, and tooling to help Cloudflare and its customers ship AI applications at internet scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer — Production Inference & Optimization
Senior ML Engineer — Production Inference & Optimization

Hc1 • Greater London

Hybrid
GBP 90,000 - 150,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Hc1 • Greater London

Hybrid
GBP 90,000 - 150,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Cloudflare • Greater London

Hybrid
GBP 110,000 - 140,000
Senior ML Engineer: Production-Scale AI & Cloud Leader
Senior ML Engineer: Production-Scale AI & Cloud Leader

Faculty Science Limited • Greater London

Hybrid
GBP 90,000 - 130,000
Unlimited Annual Leave Policy
Private healthcare and dental
Enhanced parental leave
+3
Senior ML Inference & Model Serving Engineer
Senior ML Inference & Model Serving Engineer

Google • Greater London

On-site
GBP 160,000 - 200,000
Learning opportunities
Career growth
ML Infrastructure Engineering Manager – Cloud Inference
ML Infrastructure Engineering Manager – Cloud Inference

Apple Inc. • Greater London

On-site
GBP 140,000 - 200,000
Production ML Engineer — Cloud-Native & MLOps
Production ML Engineer — Cloud-Native & MLOps

Anson McCade • Greater London

Hybrid
GBP 144,000 - 176,000
AI Engineer - Production ML Systems (Hybrid London)
AI Engineer - Production ML Systems (Hybrid London)

Formula Recruitment Limited • Greater London

Hybrid
GBP 68,000 - 108,000
Senior ML Engineer - Build & Scale Production AI Systems
Senior ML Engineer - Build & Scale Production AI Systems

Cognify Search • Greater London

On-site
GBP 90,000 - 150,000
Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions