Staff Software Engineer, GenAI Inference Performance

Google LLC

Mountain View (CA)

On-site

USD 207,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

DeepMind is seeking a Staff Software Engineer for Inference Performance Optimization, GenAI. You will work on measuring intelligence of prototypes, building systems for agent testing, and developing test problems in physics simulators, while visualizing results and leading leaderboards.

You will push AI model execution at scale, analyze the full inference stack, maximize hardware throughput, reduce cost-to-serve, and enable data-driven capacity and latency tradeoffs across teams.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.

Responsibilities

  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.

Skills

Python
C++

Education

Bachelor's degree

Job description

DeepMind is seeking a Staff Software Engineer for Inference Performance Optimization, GenAI. You will work on measuring intelligence of prototypes, building systems for agent testing, and developing test problems in physics simulators, while visualizing results and leading leaderboards.

You will push AI model execution at scale, analyze the full inference stack, maximize hardware throughput, reduce cost-to-serve, and enable data-driven capacity and latency tradeoffs across teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI, Inc. • San Francisco (CA)

On-site
USD 295,000 - 555,000
Equity
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google LLC • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff GenAI Inference Architect: High-Throughput ML Serving
Staff GenAI Inference Architect: High-Throughput ML Serving

Databricks • San Francisco (CA)

On-site
USD 190,900 - 232,800
Comprehensive benefits and perks
Annual performance bonus
Equity options
Staff AI Software Engineer: GenAI Inference on Snapdragon
Staff AI Software Engineer: GenAI Inference on Snapdragon

Qualcomm • Raleigh (NC)

On-site
USD 162,000 - 243,000
Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Databricks • San Francisco (CA)

On-site
USD 190,900 - 232,800
Comprehensive benefits and perks
Annual performance bonus
Equity options
Staff Software Engineer, AI Agents & Data Quality
Staff Software Engineer, AI Agents & Data Quality

Google Inc. • Mountain View (CA)

On-site
USD 207,000 - 300,000
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,200 - 239,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000