Senior ML Inference Engineer for LLM Serving

WeAreTechWomen

United Kingdom

On-site

GBP 120,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Google DeepMind seeks a Software Engineer to join the team responsible for bringing AI research to life, deploying large language models (LLMs) like Gemini onto production infrastructure and building robust, high-performance ML systems.

You will work across research and engineering to optimize and deploy LLMs in production, with opportunities across IC and TL tracks, across multiple teams focused on serving and inference dynamics.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining ML models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Collaborate with Research teams to understand modeling approaches with production considerations.
  • Work with infrastructure teams to deliver efficient serving infrastructure for maximum speed and scale.
  • Identify opportunities to automate tasks, remove redundancies, and improve velocity of model releases.
  • Gain deep understanding of serving frameworks, preprocessing pipelines, caching mechanisms, and related tech.
  • Leverage hardware-level profiling and roofline analysis to optimize ML frameworks, compilers (XLA), and kernels (Pallas) on accelerators.

Skills

Software development experience
ML deployment in production
Profiling ML workloads on GPUs/TPUs
Model serving infrastructure
Experience with AI systems

Education

Bachelor’s degree or equivalent practical experience

Tools

CUDA
OpenCL
JAX
PyTorch
XLA
Pallas
GPU/TPU

Job description

Google DeepMind seeks a Software Engineer to join the team responsible for bringing AI research to life, deploying large language models (LLMs) like Gemini onto production infrastructure and building robust, high-performance ML systems.

You will work across research and engineering to optimize and deploy LLMs in production, with opportunities across IC and TL tracks, across multiple teams focused on serving and inference dynamics.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer, LLM Deployment
Senior ML Systems Engineer, LLM Deployment

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Production ML Engineer - LLMs & Generative AI
Production ML Engineer - LLMs & Generative AI

Understanding Recruitment • Greater London

On-site
GBP 90,000 - 130,000
Equity
7% pension
Private healthcare
+1
ML Platform Lead - LLM Training & Inference
ML Platform Lead - LLM Training & Inference

Scale AI • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Health & Wellbeing
Career Growth stipend
Community events
+1
ML Engineer (LLMs) — Real‑World AI with Equity
ML Engineer (LLMs) — Real‑World AI with Equity

Understanding Recruitment • Greater London

On-site
GBP 90,000 - 130,000
Equity
7% pension
Private healthcare
+3
Research Engineer, Gemini Omni, DeepMind
Research Engineer, Gemini Omni, DeepMind

WeAreTechWomen • Greater London

On-site
GBP 120,000 - 170,000
Research Engineer, Gemini Omni, DeepMind
Research Engineer, Gemini Omni, DeepMind

Google DeepMind • Greater London

On-site
GBP 90,000 - 130,000
Research Scientist/Engineer, Frontier Reasoning, DeepMind
Research Scientist/Engineer, Frontier Reasoning, DeepMind

AI Chopping Block, Inc. • Greater London

Hybrid
GBP 156,000 - 226,000
Senior ML Research Engineer - Generative AI & LLMs (Hybrid)
Senior ML Research Engineer - Generative AI & LLMs (Hybrid)

Principle HR • Greater London

Hybrid
GBP 81,000 - 90,000
AI Infra Engineer: Scalable LLM Serving & Platform Design
AI Infra Engineer: Scalable LLM Serving & Platform Design

scaleai • Greater London

On-site
GBP 110,000 - 160,000