Senior ML Engineer — Production Inference & Optimization

Hc1

Greater London

Hybrid

GBP 90,000 - 150,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cloudflare is expanding its ML capability in London, building and optimizing production ML models for a serverless inference platform. You’ll collaborate with systems engineers, hardware partners, and AI/ML teams to deliver low-latency, reliable models that scale across Cloudflare’s global network.

You’ll focus on benchmarking, inference optimization, validation, and tooling to empower Workers AI and customer deployments.

Qualifications

  • Experience building production ML models with focus on reliability and performance.
  • Proficiency in Python and ML frameworks (PyTorch/TF/JAX) for model development.
  • Hands-on experience with inference optimization techniques (quantization, batching, caching).
  • Familiarity with distributed systems, networking, or serverless platforms.
  • Track record leading complex technical projects and mentoring engineers.

Responsibilities

  • Develop, optimize, and productionize ML models for Cloudflare’s serverless inference platform.
  • Build benchmarking and evaluation frameworks for latency, throughput, and cost efficiency across model families.
  • Improve inference performance via quantization, batching, caching, and runtime tuning.
  • Collaborate with systems engineers to deploy models across a heterogeneous GPU fleet and accelerators.
  • Improve model deployment workflows, validation, rollout safety, observability, and reliability.
  • Work with product/engineering teams to translate customer needs into scalable ML capabilities for Workers AI.
  • Mentor engineers and contribute to technical direction and engineering standards.

Skills

Python
ML Frameworks (PyTorch/TF/JAX)
Inference optimization
Distributed systems
Mentoring engineers

Tools

SGLang
vLLM
TensorRT-LLM
ONNX Runtime
Triton
llama.cpp

Job description

Cloudflare is expanding its ML capability in London, building and optimizing production ML models for a serverless inference platform. You’ll collaborate with systems engineers, hardware partners, and AI/ML teams to deliver low-latency, reliable models that scale across Cloudflare’s global network.

You’ll focus on benchmarking, inference optimization, validation, and tooling to empower Workers AI and customer deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: Production Inference & Optimization
Senior ML Engineer: Production Inference & Optimization

Cloudflare • Greater London

Hybrid
GBP 110,000 - 140,000
Senior ML Ops Engineer — Production AI Platforms (Hybrid)
Senior ML Ops Engineer — Production AI Platforms (Hybrid)

Boehringer Ingelheim • Greater London

Hybrid
GBP 70,000 - 110,000
Hybrid work model
Top Employer in the UK
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Hc1 • Greater London

Hybrid
GBP 90,000 - 150,000
ML Infrastructure Engineering Manager – Cloud Inference
ML Infrastructure Engineering Manager – Cloud Inference

Apple Inc. • Greater London

On-site
GBP 140,000 - 200,000
Senior ML Inference & Model Serving Engineer
Senior ML Inference & Model Serving Engineer

Google • Greater London

On-site
GBP 160,000 - 200,000
Learning opportunities
Career growth
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Cloudflare • Greater London

Hybrid
GBP 110,000 - 140,000
ML Ops Engineer: Production-Ready ML Platform
ML Ops Engineer: Production-Ready ML Platform

CMC Markets • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Cloud Engineer, ML Platforms — Hybrid Role
Senior Cloud Engineer, ML Platforms — Hybrid Role

Boehringer Ingelheim • Greater London

Hybrid
GBP 80,000 - 110,000
Production ML Engineer — Cloud-Native & MLOps
Production ML Engineer — Cloud-Native & MLOps

Anson McCade • Greater London

Hybrid
GBP 144,000 - 176,000
ML Performance Engineer: Scale & Optimize GPU/CPU Workloads
ML Performance Engineer: Scale & Optimize GPU/CPU Workloads

G-Research • Greater London

On-site
GBP 90,000 - 135,000
Lunch provided (Just Eat for Business)
Barista bar
35 days annual leave
+5