Software Engineer, Model Inference

Google

Greater London

On-site

GBP 160,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Learning opportunities
Career growth

Job summary

Google DeepMind is seeking a senior ML systems engineer to deploy and optimize large language models on Google's production infrastructure. You will work closely with research teams to design high-performance model serving solutions and drive efficiency across hardware accelerators.

The role focuses on profiling, tuning, and deploying models at scale with opportunities for IC and technical lead growth. London or Mountain View locations supported.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of software development experience.
  • 2 years deploying and maintaining ML models in production.
  • Experience profiling ML workloads on GPUs or TPUs.
  • Experience designing/optimizing model serving infrastructure.

Responsibilities

  • Collaborate with research teams to understand production-friendly modeling approaches.
  • Work with infrastructure teams to deliver efficient, high-performance serving infrastructure.
  • Identify opportunities to automate tasks, remove redundancies, and improve ML release velocity.
  • Gain deep understanding of serving frameworks, pre-processing pipelines, and caching mechanisms.
  • Use roofline analysis and hardware profiling to address bottlenecks across ML stacks.

Skills

Software development
ML model deployment
Profiling on GPUs/TPUs
Model serving infrastructure
Collaboration with researchers

Education

Bachelor’s degree or equivalent practical experience

Tools

CUDA
PyTorch
Model serving frameworks

Job description

Salary: £160,000 - 200,000 per year

Requirements:
  • We require a bachelors degree or equivalent practical experience.
  • We require 8 years of experience in software development.
  • We require 2 years of experience deploying and maintaining machine learning models in a live production environment.
  • We require experience profiling, configuring, or executing ML workloads directly on hardware accelerators such as GPUs or TPUs.
  • We require experience designing, building, or optimizing model serving infrastructure or inference backends.
Responsibilities:
  • We collaborate closely with research teams to understand next-generation modeling approaches and ensure they are designed and implemented with production considerations in mind.
  • We work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • We identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • We gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • We leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
Technologies:
  • AI
  • Hardware
  • Machine Learning
  • Model Serving
  • 3D
  • AI Agents
  • CUDA
  • LLM
  • PyTorch

More:

At Google DeepMind, we are building the worlds first general-purpose learning agent and advancing AI to solve complex global challenges and accelerate high-quality product innovation for billions of users. We work across multiple teams and offer opportunities for both individual contributor and technical lead growth, with openings for Software Engineering and Research Engineering backgrounds. This role involves bringing AI research to life by working directly with researchers and engineers to optimize and deploy large language models such as Gemini onto Googles production infrastructure. We offer learning opportunities, varied career pathways, and benefits at Google, and the position is open to preferred working locations in London, UK or Mountain View, CA, USA.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google Inc. • Greater London

On-site
GBP 150,000 - 230,000
Research Engineer, World Models, DeepMind
Research Engineer, World Models, DeepMind

Google Inc. • Greater London

On-site
GBP 120,000 - 180,000
Software Engineer, AI/ML, PhD, Early Career
Software Engineer, AI/ML, PhD, Early Career

Google Inc. • Greater London

On-site
GBP 70,000 - 90,000
Senior ML Systems Engineer — LLM Inference & Serving
Senior ML Systems Engineer — LLM Inference & Serving

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Research Engineer, World Models, DeepMind
Research Engineer, World Models, DeepMind

Google • Greater London

On-site
GBP 140,000 - 180,000
Senior ML Inference & Model Serving Engineer
Senior ML Inference & Model Serving Engineer

Google • Greater London

On-site
GBP 160,000 - 200,000
Learning opportunities
Career growth
Research Engineer, Astra, DeepMind
Research Engineer, Astra, DeepMind

Google Inc. • Greater London

On-site
GBP 120,000 - 180,000
Forward Deployed Engineer III, GCC (French, German)
Forward Deployed Engineer III, GCC (French, German)

Google Inc. • City of Westminster

On-site
GBP 80,000 - 110,000
Research Engineer, Tool Use, DeepMind
Research Engineer, Tool Use, DeepMind

Software Careers • Greater London

On-site
GBP 80,000 - 110,000