Tech Lead, ML Inference Performance

Google LLC

Greater London

On-site

GBP 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

DeepMind in London seeks a Tech Lead for Gemini Inference Performance to guide a team of performance engineers and drive optimization across ML frameworks and hardware accelerators.

You will leverage transformer and Mixture-of-Experts concepts to improve latency and memory efficiency while collaborating with research teams on deployment trade-offs. A strong leadership track record is essential.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8+ years of software development.
  • 5+ years in people management/supervision/team leadership.
  • 5+ years in a technical leadership role and overseeing projects.
  • Master’s degree or PhD in Engineering, CS, or related field preferred.
  • 5+ years in a complex, matrixed organization.

Responsibilities

  • Set the technical roadmap for a team of performance engineers and drive prioritization of optimization opportunities.
  • Identify and eliminate performance bottlenecks across ML frameworks, compilers, kernels, and serving infrastructure on hardware accelerators.
  • Apply first-principles understanding of Transformer and Mixture-of-Experts components to optimize execution and memory footprint.
  • Collaborate with research teams early to evaluate inference implications on latency and serving costs.
  • Guide development of custom kernels and serving optimizations and perform deep performance profiling.

Skills

Software development
People management
Technical leadership

Education

Bachelor's degree
Master's/PhD preferred

Tools

Transformer models
Mixture of Experts
Hardware profiling

Job description

DeepMind in London seeks a Tech Lead for Gemini Inference Performance to guide a team of performance engineers and drive optimization across ML frameworks and hardware accelerators.

You will leverage transformer and Mixture-of-Experts concepts to improve latency and memory efficiency while collaborating with research teams on deployment trade-offs. A strong leadership track record is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google LLC • Greater London

On-site
GBP 120,000 - 180,000
Model Performance Tech Lead & Manager
Model Performance Tech Lead & Manager

Google LLC • Greater London

Hybrid
GBP 198,000 - 275,000
Tech Lead, AI Inference Performance
Tech Lead, AI Inference Performance

Google • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google • Greater London

On-site
GBP 120,000 - 180,000
Senior ML Inference Engineer for LLM Serving
Senior ML Inference Engineer for LLM Serving

WeAreTechWomen • United Kingdom

On-site
GBP 120,000 - 190,000
Staff ML Performance Engineer — Edge Inference Optimizer
Staff ML Performance Engineer — Edge Inference Optimizer

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Senior GPU AI/ML Performance Engineer
Senior GPU AI/ML Performance Engineer

Google LLC • Greater London

On-site
GBP 90,000 - 150,000
Staff ML Performance Engineer: Edge Inference Optimizer
Staff ML Performance Engineer: Edge Inference Optimizer

EngineersOfAI • Greater London

Hybrid
GBP 120,000 - 180,000
Research Scientist, GenAI Safety & Alignment - Equity
Research Scientist, GenAI Safety & Alignment - Equity

Google LLC • City Of London

Hybrid
GBP 157,000 - 227,000
Tech Lead Manager, Model Performance, DeepMind
Tech Lead Manager, Model Performance, DeepMind

Google • Greater London

On-site
GBP 198,000 - 276,000