Tech Lead, Gemini Inference Performance, DeepMind

DeepMind Technologies Limited

Greater London

On-site

GBP 150,000 - 190,000

Full time

13 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Google DeepMind in London is seeking a senior performance engineering leader to set the technical roadmap for a team of engineers focused on ML framework optimization, compilers, and serving infrastructure on hardware accelerators.

You will coach reports, drive prioritization across high-leverage problems, apply transformer and Mixture-of-Experts concepts to improve latency and memory footprint, and collaborate with researchers early in development.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in a people management, supervision/team leadership role.
  • 5 years of experience in a technical leadership role and overseeing projects.

Responsibilities

  • Set the technical roadmap for a team of performance engineers, support and develop direct reports and drive prioritization across competing optimization opportunities — staying in the highest-leverage problems.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
  • Apply a first-principles understanding of Transformer and Mixture-of-Experts model components to identify and prioritize opportunities to optimize their execution efficiency and memory footprint.
  • Collaborate with research teams early in the development lifecycle to evaluate inference implications, modeling how architectural choices impact latency, memory footprint, and serving costs.
  • Guide and contribute to the development of custom kernels and serving optimizations. Conduct deep performance profiling using hardware tracing tools to analyze accelerator utilization, memory bandwidth saturation, and interconnect latency across large-scale topologies.

Skills

Software development experience
People management
Technical leadership
Roadmap planning

Education

Bachelor’s degree or equivalent practical experience
Master’s degree or PhD in Engineering or CS

Job description

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in a people management, supervision/team leadership role.
  • 5 years of experience in a technical leadership role and overseeing projects.
Minimum qualifications
  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in a people management, supervision/team leadership role.
  • 5 years of experience in a technical leadership role and overseeing projects.
Preferred qualifications
  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 5 years of experience working in a complex, matrixed organization.
About the job

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority. We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Responsibilities
  • Set the technical roadmap for a team of performance engineers, support and develop direct reports and drive prioritization across competing optimization opportunities — staying in the highest-leverage problems.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
  • Apply a first-principles understanding of Transformer and Mixture-of-Experts model components to identify and prioritize opportunities to optimize their execution efficiency and memory footprint.
  • Collaborate with research teams early in the development lifecycle to evaluate inference implications, modeling how architectural choices impact latency, memory footprint, and serving costs.
  • Guide and contribute to the development of custom kernels and serving optimizations. Conduct deep performance profiling using hardware tracing tools to analyze accelerator utilization, memory bandwidth saturation, and interconnect latency across large-scale topologies.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google LLC • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google • Greater London

On-site
GBP 120,000 - 180,000
Tech Lead Manager, Model Performance, DeepMind
Tech Lead Manager, Model Performance, DeepMind

Google • Greater London

On-site
GBP 198,000 - 276,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Tech Lead, AI Inference Performance & Optimization
Tech Lead, AI Inference Performance & Optimization

DeepMind Technologies Limited • Greater London

On-site
GBP 150,000 - 190,000
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google DeepMind • Greater London

On-site
GBP 140,000 - 230,000
Tech Lead, AI Inference Performance
Tech Lead, AI Inference Performance

Google • Greater London

On-site
GBP 120,000 - 180,000
Research Engineer, Gemini Omni, DeepMind
Research Engineer, Gemini Omni, DeepMind

Google DeepMind • Greater London

On-site
GBP 90,000 - 130,000
Research Engineer, Gemini Omni, DeepMind
Research Engineer, Gemini Omni, DeepMind

Google • Greater London

On-site
GBP 120,000 - 170,000
Research Scientist, Gemini Safety and Behavior, DeepMind
Research Scientist, Gemini Safety and Behavior, DeepMind

DeepMind Technologies Limited • Greater London

On-site
GBP 120,000 - 180,000
Equity
Benefits