Staff AI Inference Performance Engineer (GenAI)

Google

Town of Montana (WI)

On-site

USD 207,000 - 300,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus target
Equity

Job summary

Google DeepMind is hiring a Staff Software Engineer to optimize AI inference performance at scale. You will analyze the entire inference stack, implement optimization techniques, and drive systemic improvements to maximize throughput and reduce cost per inference.

You will collaborate with ML and systems teams on a range of AI agents and research prototypes. The role emphasizes end-to-end performance engineering, profiling workloads, and building instrumentation to guide capacity and latency

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of software development experience.
  • Experience with Python and C++, including navigating and debugging serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth, and modern serving architectures.

Responsibilities

  • Analyze and optimize AI inference workloads to increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference bottlenecks across the stack.
  • Model latency-to-cost impacts and translate insights into production signals.
  • Develop tools and metrics to track compute spend across the fleet.

Skills

Python
C++
Serving codebases
Performance optimization

Education

Bachelor's degree in CS/CE/EE or related field

Tools

vLLM
TensorRT-LLM
SGLang
Dynamo

Job description

Google DeepMind is hiring a Staff Software Engineer to optimize AI inference performance at scale. You will analyze the entire inference stack, implement optimization techniques, and drive systemic improvements to maximize throughput and reduce cost per inference.

You will collaborate with ML and systems teams on a range of AI agents and research prototypes. The role emphasizes end-to-end performance engineering, profiling workloads, and building instrumentation to guide capacity and latency

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GenAI Inference Performance Engineer
Staff GenAI Inference Performance Engineer

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Bonus target
Equity
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
AI Inference Compute Engineer — Hardware/Software Co-Design
AI Inference Compute Engineer — Hardware/Software Co-Design

Google DeepMind • San Francisco (CA)

On-site
USD 174,000 - 252,000
Staff Software Engineer, Applied AI & GenAI Platforms
Staff Software Engineer, Applied AI & GenAI Platforms

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Staff Software Engineer, AI Agents & Data Quality
Staff Software Engineer, AI Agents & Data Quality

Google Inc. • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff GenAI Inference Architect: High-Throughput ML Serving
Staff GenAI Inference Architect: High-Throughput ML Serving

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Senior AI/ML Systems Architect
Senior AI/ML Systems Architect

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff AI Engineer — LLM & Agent Evaluation
Staff AI Engineer — LLM & Agent Evaluation

Socket.dev • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff GenAI Inference Engineer: Optimize LLM Serving Latency
Staff GenAI Inference Engineer: Optimize LLM Serving Latency

Menlo Ventures • San Francisco (CA)

On-site
USD 190,000 - 233,000
Annual performance bonus
Equity options
Comprehensive health benefits