Senior ML Inference & Serving Engineer (LLM Deploy)

Google DeepMind

Mountain View (CA)

On-site

USD 207,000 - 300,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google DeepMind in Mountain View, CA is seeking a Software Engineer to advance AI agents and deploy large language models on production infrastructure. You will collaborate with researchers and engineers to optimize serving frameworks, test new algorithms, and scale inference on GPUs/TPUs.

The role offers IC and TL opportunities across teams, with a strong emphasis on performance, reliability, and impact. The team culture emphasizes collaboration, experimentation, and measurable progress in AI

Qualifications

  • Bachelor's degree or equivalent practical experience required.
  • 8 years of software development experience.
  • 2 years deploying and maintaining ML models in production.
  • Experience with hardware accelerators (GPU/TPU).
  • Experience designing or optimizing model serving infrastructure.

Responsibilities

  • Collaborate with Research teams to design production-friendly modeling approaches.
  • Work with infrastructure teams to deliver efficient serving infrastructure.
  • Identify opportunities to automate tasks and improve model release velocity.
  • Understand serving frameworks, pre-processing pipelines, caching, and related tech.
  • Use profiling and systems analysis to eliminate bottlenecks across ML frameworks and kernels.

Skills

Bachelor's degree or equivalent
8 years experience in software dev
2 years deploying ML models in prod
Hardware accelerators (GPU/TPU) exp
Model serving infra design

Education

Bachelor's degree or equivalent practical experience

Tools

Serving infra development
ML frameworks exposure (JAX, PyTorch)

Job description

Google DeepMind in Mountain View, CA is seeking a Software Engineer to advance AI agents and deploy large language models on production infrastructure. You will collaborate with researchers and engineers to optimize serving frameworks, test new algorithms, and scale inference on GPUs/TPUs.

The role offers IC and TL opportunities across teams, with a strong emphasis on performance, reliability, and impact. The team culture emphasizes collaboration, experimentation, and measurable progress in AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Research Scientist — Scalable ML & LLM Systems
Senior AI Research Scientist — Scalable ML & LLM Systems

Google DeepMind • Mountain View (CA)

On-site
USD 262,000 - 364,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 207,000 - 300,000
Applied ML Research Engineer - LLMs to Production
Applied ML Research Engineer - LLMs to Production

Google DeepMind • San Francisco (CA)

On-site
USD 174,000 - 252,000
Equity
Benefits package
Senior AI/ML Systems Engineer: LLM Inference on GDC
Senior AI/ML Systems Engineer: LLM Inference on GDC

Socket.dev • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Senior Staff AI Inference Platform Architect LLM, Agents
Senior Staff AI Inference Platform Architect LLM, Agents

Google • Kirkland (WA)

On-site
USD 262,000 - 365,000
Health insurance
Retirement Benefits
Paid Time Off
+4
Senior ML Infra Engineer for Scalable LLM Systems
Senior ML Infra Engineer for Scalable LLM Systems

Moveworks • Mountain View (CA)

On-site
USD 120,000 - 160,000
Senior AI Research Scientist — Large-Scale ML & LLM Systems
Senior AI Research Scientist — Large-Scale ML & LLM Systems

Socket.dev • Mountain View (CA)

On-site
USD 262,000 - 364,000
25% bonus target
Equity
Benefits
Research Engineer: LLMs, ML Systems & Production
Research Engineer: LLMs, ML Systems & Production

Socket.dev • Mountain View (CA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Senior Software Engineer, AI/LLM on GDC Platform
Senior Software Engineer, AI/LLM on GDC Platform

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000