ML Inference & Serving Engineer

Google Inc.

Greater London

Hybrid

GBP 153,000 - 222,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google DeepMind in London, UK and Mountain View, CA is seeking a Software Engineer for model inference. You will work with researchers to optimize and deploy large language models (LLMs) on Google's production infrastructure, building scalable serving backends and performance-driven tests.

The role welcomes both IC and TL paths and offers opportunities across multiple teams. A strong software engineering foundation and experience with ML workloads on accelerators is essential.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining machine learning models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Create systems for agent testing in 2D and 3D games and develop test problems within physics simulators.
  • Develop graphical visualizations of results and build competitive agent leaderboards.
  • Test new algorithms on robots and collaborate with ML/ neuroscience teams.
  • Optimize and deploy large language models (LLMs) like Gemini on production infrastructure.

Skills

Software development
ML deployment
Profiling on GPUs/TPUs
Model serving infrastructure
Systems optimization

Education

Bachelor's degree or equivalent practical experience

Tools

JAX
PyTorch
CUDA
OpenCL
Pallas
XLA

Job description

Google DeepMind in London, UK and Mountain View, CA is seeking a Software Engineer for model inference. You will work with researchers to optimize and deploy large language models (LLMs) on Google's production infrastructure, building scalable serving backends and performance-driven tests.

The role welcomes both IC and TL paths and offers opportunities across multiple teams. A strong software engineering foundation and experience with ML workloads on accelerators is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer — LLM Inference & Serving
Senior ML Systems Engineer — LLM Inference & Serving

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

United States Digital Space LLC • Greater London

Hybrid
GBP 133,000 - 222,000
Bonus target
Equity
Benefits
Senior ML Systems Engineer - LLM Deployment & Infra
Senior ML Systems Engineer - LLM Deployment & Infra

United States Digital Space LLC • Greater London

Hybrid
GBP 133,000 - 222,000
Bonus target
Equity
Benefits
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google Inc. • Greater London

Hybrid
GBP 153,000 - 222,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

G-Research • Greater London

On-site
GBP 90,000 - 150,000
Discretionary bonus
35 days leave
Pension contributions
+2
ML Infrastructure Engineer: Scalable LLM Serving Platform
ML Infrastructure Engineer: Scalable LLM Serving Platform

Neura Market • Greater London

On-site
GBP 110,000 - 165,000
AI Infrastructure Engineer, Serving Platform Scale AI London, UK
AI Infrastructure Engineer, Serving Platform Scale AI London, UK

Neura Market • Greater London

On-site
GBP 110,000 - 165,000
Researcher, Training - London
Researcher, Training - London

United States Digital Space LLC • Greater London

Hybrid
GBP 70,000 - 90,000
Relocation support
Hybrid work schedule
AI Infrastructure Engineer, Serving Platform
AI Infrastructure Engineer, Serving Platform

Scale AI • Greater London

On-site
GBP 90,000 - 130,000