Senior ML Inference & Model Serving Engineer

Google

Greater London

On-site

GBP 160,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Learning opportunities
Career growth

Job summary

Google DeepMind is seeking a senior ML systems engineer to deploy and optimize large language models on Google's production infrastructure. You will work closely with research teams to design high-performance model serving solutions and drive efficiency across hardware accelerators.

The role focuses on profiling, tuning, and deploying models at scale with opportunities for IC and technical lead growth. London or Mountain View locations supported.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of software development experience.
  • 2 years deploying and maintaining ML models in production.
  • Experience profiling ML workloads on GPUs or TPUs.
  • Experience designing/optimizing model serving infrastructure.

Responsibilities

  • Collaborate with research teams to understand production-friendly modeling approaches.
  • Work with infrastructure teams to deliver efficient, high-performance serving infrastructure.
  • Identify opportunities to automate tasks, remove redundancies, and improve ML release velocity.
  • Gain deep understanding of serving frameworks, pre-processing pipelines, and caching mechanisms.
  • Use roofline analysis and hardware profiling to address bottlenecks across ML stacks.

Skills

Software development
ML model deployment
Profiling on GPUs/TPUs
Model serving infrastructure
Collaboration with researchers

Education

Bachelor’s degree or equivalent practical experience

Tools

CUDA
PyTorch
Model serving frameworks

Job description

Google DeepMind is seeking a senior ML systems engineer to deploy and optimize large language models on Google's production infrastructure. You will work closely with research teams to design high-performance model serving solutions and drive efficiency across hardware accelerators.

The role focuses on profiling, tuning, and deploying models at scale with opportunities for IC and technical lead growth. London or Mountain View locations supported.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer — LLM Inference & Serving
Senior ML Systems Engineer — LLM Inference & Serving

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
Software Engineer, Model Inference
Software Engineer, Model Inference

Google • Greater London

On-site
GBP 160,000 - 200,000
Learning opportunities
Career growth
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Greater London

Hybrid
GBP 153,000 - 222,000
Equity
Bonus target 20%
Comprehensive benefits
SRE Systems Engineering Manager, ML Compute
SRE Systems Engineering Manager, ML Compute

Google Inc. • Greater London

Hybrid
GBP 120,000 - 180,000
Lead, Generative AI Engineering & ML Infra
Lead, Generative AI Engineering & ML Infra

Google Inc. • Greater London

On-site
GBP 150,000 - 230,000
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google Inc. • Greater London

On-site
GBP 150,000 - 230,000
Senior ML Engineer: Production Inference & Optimization
Senior ML Engineer: Production Inference & Optimization

Cloudflare • Greater London

Hybrid
GBP 110,000 - 140,000
Research Engineer, World Models, DeepMind
Research Engineer, World Models, DeepMind

Google Inc. • Greater London

On-site
GBP 120,000 - 180,000
Research Manager: Production Inference Systems
Research Manager: Production Inference Systems

DeepL • Greater London

Hybrid
GBP 90,000 - 130,000
Diverse and internationally dispersed팀
Hybrid work
Virtual Shares
+4
Lead ML Infra Engineer for Scalable Physics AI
Lead ML Infra Engineer for Scalable Physics AI

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000