Software Engineer, Model Inference & LLM Deployment

Google LLC

Mountain View, Northern (CA, KY)

Hybrid

USD 180,000 - 300,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Google DeepMind seeks a Software Engineer for Model Inference to advance production-grade ML serving systems. You will collaborate with ML researchers and engineers to deploy large language models and optimize inference on GPUs/TPUs within Google's infra.

You'll work across model hosting, testing, and performance tuning, contributing to scalable, low-latency serving infrastructure and end-to-end deployment pipelines in a mission-driven team.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining ML models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Responsibilities

  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Skills

ML deployment
Performance tuning
Collaboration
Software engineering

Education

Bachelor’s degree or equivalent practical experience
8 years of software development experience
2 years deploying ML models in production
Experience with ML workloads on GPUs/TPUs
Model serving/inference backend design

Tools

JAX
PyTorch
CUDA/OpenCL
XLA
Pallas

Job description

Google DeepMind seeks a Software Engineer for Model Inference to advance production-grade ML serving systems. You will collaborate with ML researchers and engineers to deploy large language models and optimize inference on GPUs/TPUs within Google's infra.

You'll work across model hosting, testing, and performance tuning, contributing to scalable, low-latency serving infrastructure and end-to-end deployment pipelines in a mission-driven team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Engineer - LLM Deployment & Serving
ML Inference Engineer - LLM Deployment & Serving

Google DeepMind • Mountain View (CA)

Hybrid
USD 230,000 - 290,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Mountain View (CA)

Hybrid
USD 230,000 - 290,000
Senior ML Engineer – AI/NLP & RL Systems
Senior ML Engineer – AI/NLP & RL Systems

Google • Mountain View (CA)

On-site
USD 174,000 - 252,000
Bonus target
Equity grant
Benefits
Senior Software Engineer, Distributed Cloud AI & LLM
Senior Software Engineer, Distributed Cloud AI & LLM

Google • Town of Montana (WI)

On-site
USD 174,000 - 252,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google LLC • Mountain View (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Senior AI/ML Engineer — Build Scalable Models
Senior AI/ML Engineer — Build Scalable Models

Google • Mountain View (CA)

On-site
USD 174,000 - 252,000
GenAI Research Engineer — LLM & Model Evaluation
GenAI Research Engineer — LLM & Model Evaluation

Google Inc. • Cambridge (MA), Northern (KY)

Hybrid
USD 174,000 - 252,000
Staff AI/ML Data & Model Infra Engineer
Staff AI/ML Data & Model Infra Engineer

Google • United States

On-site
USD 207,000 - 300,000
Senior AI/ML Research Engineer - LLMs & Production Systems
Senior AI/ML Research Engineer - LLMs & Production Systems

Google • Mountain View (CA)

On-site
USD 174,000 - 252,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000