ML Inference Engineer - LLM Deployment & Serving

Google DeepMind

Mountain View (CA)

Hybrid

USD 230,000 - 290,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Google DeepMind seeks a Software Engineer to advance large-scale ML deployment and research-to-production workflows. You will collaborate with ML researchers to optimize model serving on GPUs/TPUs, build efficient test infrastructure, and deploy agents and LLMs on Google's production stack.

You'll work across IC and TL paths, contributing to performance tuning, architecture decisions, and end-to-end lifecycle tooling for research-to-production pipelines in a mission-driven team.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining machine learning models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Skills

Software development
ML deployment
Hardware accelerators
Model serving

Education

Bachelor’s degree

Tools

JAX
PyTorch
CUDA/OpenCL

Job description

Google DeepMind seeks a Software Engineer to advance large-scale ML deployment and research-to-production workflows. You will collaborate with ML researchers to optimize model serving on GPUs/TPUs, build efficient test infrastructure, and deploy agents and LLMs on Google's production stack.

You'll work across IC and TL paths, contributing to performance tuning, architecture decisions, and end-to-end lifecycle tooling for research-to-production pipelines in a mission-driven team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Inference Engineer: ML Serving & Hardware
Model Inference Engineer: ML Serving & Hardware

Google Inc. • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior AI/ML Engineer - LLMs, RL & Production ML
Senior AI/ML Engineer - LLMs, RL & Production ML

Socket.dev • Mountain View (CA)

On-site
USD 174,000 - 252,000
15% bonus target
Equity
Benefits
Senior Software Engineer, Distributed Cloud AI & LLM
Senior Software Engineer, Distributed Cloud AI & LLM

Google • Town of Montana (WI)

On-site
USD 174,000 - 252,000
Senior AI/ML Research Engineer - LLMs & Production Systems
Senior AI/ML Research Engineer - LLMs & Production Systems

Google • Mountain View (CA)

On-site
USD 174,000 - 252,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Mountain View (CA)

Hybrid
USD 230,000 - 290,000
Staff ML Frameworks Engineer – Scalable ML Infra
Staff ML Frameworks Engineer – Scalable ML Infra

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff AI Engineer — LLM & Agent Evaluation
Staff AI Engineer — LLM & Agent Evaluation

Socket.dev • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff ML Systems Engineer - LLM Serving & RL
Staff ML Systems Engineer - LLM Serving & RL

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
ML Infra Engineer: Scale RL Pipelines & Inference
ML Infra Engineer: Scale RL Pipelines & Inference

Moonfire • Paris (TX)

On-site
USD 115,000 - 173,000
Equity
Flexible time off
Relocation package
+4
Lead ML Architect for High-Scale Ads & AI
Lead ML Architect for High-Scale Ads & AI

Google • Pittsburgh

On-site
USD 207,000 - 300,000