Senior ML Systems Engineer — LLM Inference & Serving

Google DeepMind

Greater London

Hybrid

GBP 153,000 - 222,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Bonus target 20%
Comprehensive benefits

Job summary

Google DeepMind seeks a Software Engineer to advance AI agents and deploy large language models on production infrastructure. You will build testing systems, visualize results, and optimize inference paths across GPUs/TPUs.

Collaborate with researchers and engineers, contributing to scalable serving backends and ensuring performance for a broad user base.

This role offers opportunities for IC and TL tracks across multiple teams, with emphasis on efficient, safe, and ethical AI deployment.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining machine learning models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Skills

Software development
ML deployment in production
Profiling on accelerators
Model serving infrastructure

Education

Bachelor's degree or equivalent practical experience

Tools

JAX
PyTorch
CUDA
OpenCL
Pallas
XLA
GPUs/TPUs

Job description

Google DeepMind seeks a Software Engineer to advance AI agents and deploy large language models on production infrastructure. You will build testing systems, visualize results, and optimize inference paths across GPUs/TPUs.

Collaborate with researchers and engineers, contributing to scalable serving backends and ensuring performance for a broad user base.

This role offers opportunities for IC and TL tracks across multiple teams, with emphasis on efficient, safe, and ethical AI deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference & Serving Engineer
ML Inference & Serving Engineer

Google Inc. • Greater London

Hybrid
GBP 153,000 - 222,000
Senior ML Systems Engineer - LLM Deployment & Infra
Senior ML Systems Engineer - LLM Deployment & Infra

United States Digital Space LLC • Greater London

Hybrid
GBP 133,000 - 222,000
Bonus target
Equity
Benefits
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

United States Digital Space LLC • Greater London

Hybrid
GBP 133,000 - 222,000
Bonus target
Equity
Benefits
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google Inc. • Greater London

Hybrid
GBP 153,000 - 222,000
Staff Software Engineer, GenAI & AI Systems Architect
Staff Software Engineer, GenAI & AI Systems Architect

Google Inc. • City of Westminster

On-site
GBP 120,000 - 180,000
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google Inc. • Greater London

On-site
GBP 150,000 - 230,000
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google DeepMind • Greater London

On-site
GBP 150,000 - 210,000
Applied AI Engineering Manager: Generative ML Systems
Applied AI Engineering Manager: Generative ML Systems

Google DeepMind • Greater London

On-site
GBP 150,000 - 210,000
Senior AI Engineer — Production-Grade LLM Systems
Senior AI Engineer — Production-Grade LLM Systems

Indicium AI • Greater London

On-site
GBP 90,000 - 130,000
Fast-growing start-up
Highly competitive salary package with
Company bonus
+3