Senior ML Systems Engineer, LLM Deployment

DeepMind Technologies Limited

Greater London

Hybrid

GBP 140,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Google DeepMind is seeking a Software Engineer to advance the deployment of large language models onto production infrastructure. You will collaborate with researchers and engineers to optimize serving, test agent systems, and build scalable inference backends.

The role involves working with cutting-edge AI, integrating ML workloads on GPUs/TPUs, and contributing to end-to-end AI deployment across multiple teams and projects.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of software development experience.
  • 2 years deploying and maintaining ML models in a live production environment.
  • Experience with profiling, configuring, or executing ML workloads on hardware accelerators (GPU/TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Skills

Software development
ML deployment
Model serving
Hardware accelerators

Education

Bachelor’s degree or equivalent practical experience

Tools

JAX
PyTorch
CUDA
OpenCL
XLA
Pallas

Job description

Google DeepMind is seeking a Software Engineer to advance the deployment of large language models onto production infrastructure. You will collaborate with researchers and engineers to optimize serving, test agent systems, and build scalable inference backends.

The role involves working with cutting-edge AI, integrating ML workloads on GPUs/TPUs, and contributing to end-to-end AI deployment across multiple teams and projects.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer for LLM Serving
Senior ML Inference Engineer for LLM Serving

WeAreTechWomen • United Kingdom

On-site
GBP 120,000 - 190,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Lead Architect — Generative AI & ML Infrastructure
Lead Architect — Generative AI & ML Infrastructure

Google DeepMind • Greater London

On-site
GBP 140,000 - 230,000
Production ML Engineer - LLMs & Generative AI
Production ML Engineer - LLMs & Generative AI

Understanding Recruitment • Greater London

On-site
GBP 90,000 - 130,000
Equity
7% pension
Private healthcare
+1
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google DeepMind • Greater London

On-site
GBP 140,000 - 230,000
Senior AI Engineer: Build Scalable LLM Solutions
Senior AI Engineer: Build Scalable LLM Solutions

Addition • Greater London

On-site
GBP 20,295,000 - 24,908,000
ML Platform Lead - LLM Training & Inference
ML Platform Lead - LLM Training & Inference

Scale AI • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Health & Wellbeing
Career Growth stipend
Community events
+1
Production AI Engineer: Real-World LLM Systems
Production AI Engineer: Real-World LLM Systems

SLR Consulting • Greater London

Hybrid
GBP 70,000 - 110,000
Senior ML Platform Engineer - Scalable AI Deployment
Senior ML Platform Engineer - Scalable AI Deployment

Scale AI, Inc. • Greater London

On-site
GBP 100,000 - 150,000
Tech Lead Manager, Model Performance, DeepMind
Tech Lead Manager, Model Performance, DeepMind

WeAreTechWomen • Greater London

Hybrid
GBP 198,000 - 276,000