Tech Lead, Gemini Inference Performance, DeepMind

Google

Greater London

On-site

GBP 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Google is seeking a senior leader to set the technical roadmap for a team of performance engineers, focusing on optimizing ML frameworks, kernels, and serving infrastructure. You will collaborate with researchers to balance latency, memory, and cost, guiding impactful optimizations across large-scale systems.

The role requires substantial experience in software leadership, performance analysis, and deep knowledge of Transformer/MoE architectures, with a push to deliver high-impact improvements

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years in people management/leadership roles.
  • 5 years in technical leadership overseeing projects.

Responsibilities

  • Set the technical roadmap for a team of performance engineers and drive prioritization of optimization opportunities.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify bottlenecks across ML frameworks and serving infrastructure.
  • Apply first-principles understanding of Transformer and Mixture-of-Experts components to optimize execution efficiency and memory footprint.
  • Collaborate with research teams to evaluate inference implications and model how architecture affects latency and serving costs.
  • Guide development of custom kernels and serving optimizations; conduct deep profiling of accelerator utilization and memory bandwidth.

Skills

Performance engineering
Team leadership
ML frameworks
Profiling tools
Transformer
Mixture of Experts

Education

Bachelor's degree
Master's or PhD in Engineering/CS

Tools

Roofline analysis
Hardware profiling tools
ML frameworks
Custom kernels

Job description

Minimum qualifications:
  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in a people management, supervision/team leadership role.
  • 5 years of experience in a technical leadership role and overseeing projects.

Preferred qualifications:
  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 5 years of experience working in a complex, matrixed organization.
About the job

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Responsibilities
  • Set the technical roadmap for a team of performance engineers, support and develop direct reports and drive prioritization across competing optimization opportunities — staying in the highest-leverage problems.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
  • Apply a first-principles understanding of Transformer and Mixture-of-Experts model components to identify and prioritize opportunities to optimize their execution efficiency and memory footprint.
  • Collaborate with research teams early in the development lifecycle to evaluate inference implications, modeling how architectural choices impact latency, memory footprint, and serving costs.
  • Guide and contribute to the development of custom kernels and serving optimizations. Conduct deep performance profiling using hardware tracing tools to analyze accelerator utilization, memory bandwidth saturation, and interconnect latency across large-scale topologies.

Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, Gemini Inference Performance, DeepMind
Tech Lead, Gemini Inference Performance, DeepMind

Google LLC • Greater London

On-site
GBP 120,000 - 180,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Tech Lead Manager, Model Performance, DeepMind
Tech Lead Manager, Model Performance, DeepMind

Google LLC • Greater London

Hybrid
GBP 198,000 - 275,000
Tech Lead Manager, Model Performance, DeepMind
Tech Lead Manager, Model Performance, DeepMind

Google • Greater London

On-site
GBP 198,000 - 276,000
Senior GPU Performance Software Engineer, AI/ML
Senior GPU Performance Software Engineer, AI/ML

Google LLC • Greater London

On-site
GBP 90,000 - 150,000
Manager, Applied AI Engineering, DeepMind
Manager, Applied AI Engineering, DeepMind

Google Inc. • Greater London

On-site
GBP 150,000 - 210,000
Senior Program Manager, Frontier Strategy and Governance, DeepMind Institute, DeepMind
Senior Program Manager, Frontier Strategy and Governance, DeepMind Institute, DeepMind

Google LLC • Greater London

On-site
GBP 110,000 - 150,000
Research Engineer, Gemini Omni, DeepMind
Research Engineer, Gemini Omni, DeepMind

Google DeepMind • Greater London

On-site
GBP 90,000 - 130,000
Research Engineer, Gemini Omni, DeepMind
Research Engineer, Gemini Omni, DeepMind

Google • Greater London

On-site
GBP 120,000 - 170,000
Research Scientist, Gemini Safety and Behavior, DeepMind
Research Scientist, Gemini Safety and Behavior, DeepMind

Google LLC • City Of London

Hybrid
GBP 157,000 - 227,000