Software Engineer, Model Inference, DeepMind

Google

United States

On-site

USD 207,000 - 300,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google DeepMind is hiring a Software Engineer to advance model inference and deployment in production environments, designing scalable serving backends and optimization techniques.

You will collaborate with researchers and engineers to bring AI models to life, deploy LLMs like Gemini, and ensure high performance on Google’s production infrastructure across GPUs/TPUs.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining ML models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Skills

Software development
ML deployment
Performance profiling

Education

Bachelor’s degree or equivalent

Tools

JAX
PyTorch
CUDA
OpenCL
XLA

Job description

Software Engineer, Model Inference, DeepMind

Location: London, UK; Mountain View, CA, USA

Minimum qualifications:
  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining machine learning models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.
Preferred qualifications:
  • Experience with developing serving infrastructure.
  • Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL).
  • Experience profiling software to identify performance bottlenecks.
  • Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism).
  • Familiarity with writing performance-optimized kernels.
  • Understanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling).
About the job

At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.

In this role, you will be at the forefront of bringing AI research to life. You'll work directly with researchers and engineers to optimize and deploy large language models (LLMs) like Gemini onto Google's production infrastructure, impacting users across a different range of applications. This involves a blend of technical expertise and collaborative problem-solving to ensure both efficiency and quality throughout the entire LLM deployment lifecycle.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities
  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 207,000 - 300,000
Software Engineer, Science and Strategic Initiatives, DeepMind
Software Engineer, Science and Strategic Initiatives, DeepMind

Google • United States

Hybrid
USD 174,000 - 252,000
Software Technical Lead, On-Device Inference Software, Robotics, DeepMind
Software Technical Lead, On-Device Inference Software, Robotics, DeepMind

Google • Mountain View (CA)

On-site
USD 207,000 - 301,000
Staff Software Engineer, Performance and Kernel, DeepMind
Staff Software Engineer, Performance and Kernel, DeepMind

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, Data Pipelines, DeepMind
Staff Software Engineer, Data Pipelines, DeepMind

Google DeepMind • New York (NY)

On-site
USD 207,000 - 300,000
Staff Research Engineer, Applied AI, DeepMind
Staff Research Engineer, Applied AI, DeepMind

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Technical Program Manager, Model Launches, AI Studio, DeepMind
Technical Program Manager, Model Launches, AI Studio, DeepMind

Google • Mountain View (CA)

On-site
USD 256,000 - 279,000
Research Engineer, Gemini Information Task, DeepMind
Research Engineer, Gemini Information Task, DeepMind

Google • United States

On-site
USD 174,000 - 252,000
Equity
Bonus target
Software Engineer, Science and Strategic Initiatives, DeepMind
Software Engineer, Science and Strategic Initiatives, DeepMind

Google Inc. • Mountain View (CA)

On-site
USD 174,000 - 252,000
Forward Deployed Engineer, DeepMind
Forward Deployed Engineer, DeepMind

Google • United States

On-site
USD 174,000 - 252,000