Tech Lead, AI Inference Performance & Optimization

DeepMind Technologies Limited

Greater London

On-site

GBP 150,000 - 190,000

Full time

9 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Google DeepMind in London is seeking a senior performance engineering leader to set the technical roadmap for a team of engineers focused on ML framework optimization, compilers, and serving infrastructure on hardware accelerators.

You will coach reports, drive prioritization across high-leverage problems, apply transformer and Mixture-of-Experts concepts to improve latency and memory footprint, and collaborate with researchers early in development.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in a people management, supervision/team leadership role.
  • 5 years of experience in a technical leadership role and overseeing projects.

Responsibilities

  • Set the technical roadmap for a team of performance engineers, support and develop direct reports and drive prioritization across competing optimization opportunities — staying in the highest-leverage problems.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
  • Apply a first-principles understanding of Transformer and Mixture-of-Experts model components to identify and prioritize opportunities to optimize their execution efficiency and memory footprint.
  • Collaborate with research teams early in the development lifecycle to evaluate inference implications, modeling how architectural choices impact latency, memory footprint, and serving costs.
  • Guide and contribute to the development of custom kernels and serving optimizations. Conduct deep performance profiling using hardware tracing tools to analyze accelerator utilization, memory bandwidth saturation, and interconnect latency across large-scale topologies.

Skills

Software development experience
People management
Technical leadership
Roadmap planning

Education

Bachelor’s degree or equivalent practical experience
Master’s degree or PhD in Engineering or CS

Job description

Google DeepMind in London is seeking a senior performance engineering leader to set the technical roadmap for a team of engineers focused on ML framework optimization, compilers, and serving infrastructure on hardware accelerators.

You will coach reports, drive prioritization across high-leverage problems, apply transformer and Mixture-of-Experts concepts to improve latency and memory footprint, and collaborate with researchers early in development.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, AI Inference Performance
Tech Lead, AI Inference Performance

Google • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead, ML Inference Performance
Tech Lead, ML Inference Performance

Google LLC • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

DeepMind Technologies Limited • Greater London

On-site
GBP 150,000 - 190,000
Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google LLC • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead Manager, Model Performance, DeepMind
Tech Lead Manager, Model Performance, DeepMind

Google • Greater London

On-site
GBP 198,000 - 276,000
Senior GPU AI/ML Performance Engineer
Senior GPU AI/ML Performance Engineer

Google LLC • Greater London

On-site
GBP 90,000 - 150,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Senior Performance Engineer — AI Inference & Systems Optimizer
Senior Performance Engineer — AI Inference & Systems Optimizer

CommonAI CIC • Cambridge

On-site
GBP 70,000 - 100,000
Competitive salary
Pension
Professional development
+2
AI Inference Optimizations Engineer (MLIR/LLVM)
AI Inference Optimizations Engineer (MLIR/LLVM)

microTECH Global Limited • City of Edinburgh

Hybrid
GBP 90,000 - 120,000