Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google

Mountain View (CA)

On-site

USD 207,000 - 300,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Bonus target
Benefits

Job summary

Google DeepMind is seeking a Staff Software Engineer focused on Inference Performance Optimization for GenAI. Based in Mountain View, you will drive optimization of AI inference workloads, design fast serving techniques, and analyze bottlenecks to maximize throughput and minimize latency across distributed systems.

You will work with a mission-driven team advancing AI agents, with emphasis on scalable inference, profiling, and performance engineering.

Qualifications

  • Bachelor's degree or equivalent practical experience in CS/CE/Applied Math or related field.
  • 8 years of software development experience.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints and modern serving architectures.

Responsibilities

  • Analyze and optimize AI inference workloads to increase throughput per GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model latency-to-cost impacts and translate insights into actionable production signals.
  • Develop investigative tools and metrics to track compute usage across the fleet.

Skills

Python
C++
Debugging
Serving architectures

Education

Bachelor's degree or equivalent experience

Job description

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

DeepMind – Mountain View, CA, USA

Minimum qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.
Preferred qualifications:
  • Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).
  • Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.
  • Experience with observability and reliability for large distributed systems.
  • Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.
About the job

At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.

As an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities
  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.

Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google • United States

On-site
USD 207,000 - 300,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 207,000 - 300,000
Software Technical Lead, On-Device Inference Software, Robotics, DeepMind
Software Technical Lead, On-Device Inference Software, Robotics, DeepMind

Google • Mountain View (CA)

On-site
USD 207,000 - 301,000
Engineering Manager, DSC AI Inference Platform
Engineering Manager, DSC AI Inference Platform

Google • Sunnyvale (CA)

On-site
USD 207,000 - 301,000
Senior Staff Engineer, GDC AI Inference Platform
Senior Staff Engineer, GDC AI Inference Platform

Google • United States

On-site
USD 262,000 - 365,000
Health insurance
401(k) match
Paid time off
+4
Staff Research Engineer, Applied AI, DeepMind
Staff Research Engineer, Applied AI, DeepMind

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, AI/ML GenAI, Google Cloud AI
Staff Software Engineer, AI/ML GenAI, Google Cloud AI

Google • United States

On-site
USD 207,000 - 301,000
Health, dental, vision, life, disabled
401(k) with company match
Paid time off: 20 days/year
+4
Tech Lead Manager, Staff Software Engineering, XProf
Tech Lead Manager, Staff Software Engineering, XProf

Google • Sunnyvale (CA)

On-site
USD 207,000 - 301,000
Research Engineer, DeepMind
Research Engineer, DeepMind

Google • United States

On-site
USD 262,000 - 365,000
Equity
Benefits
Software Engineer, Science and Strategic Initiatives, DeepMind
Software Engineer, Science and Strategic Initiatives, DeepMind

Google • United States

Hybrid
USD 174,000 - 252,000