Edge & Hybrid Cloud AI Engineer - LLM Inference Lead

Google

Sunnyvale (CA)

On-site

USD 174,000 - 252,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits
Bonus target

Job summary

Google is hiring software engineers to advance AI at scale, delivering LLM inference serving within the Google Distributed Cloud (GDC) Platform. You will work on critical components, from model lifecycle to efficient data loading and dynamic routing, enabling AI across hybrid cloud environments.

You will collaborate with teams across core platform services, tackle complex systems challenges, and contribute to a scalable serving infrastructure that supports innovative AI workloads.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience with software development in one or more programming languages.
  • 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
  • 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
  • Experience programming in Go for software development, including AI/ML applications.

Responsibilities

  • Lead the technical design, development, and optimization of software components critical for Large Language Model (LLM) inference serving on GDC. This includes model life-cycle management, efficient data loading, dynamic request routing, and intelligent load balancing.
  • Drive horizontal integration across core platform services, including billing, logging, observability, security, and quota management.
  • Implement and enhance serving capabilities to support advanced LLM techniques like disaggregated serving, speculative decoding, quantization, and efficient model sharding across distributed hardware.
  • Collaborate closely with internal teams developing core LLM frameworks, container orchestration (Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure, and hardware acceleration to build a cohesive and high-performance serving platform.

Skills

Go programming
AI/ML applications
Distributed systems
Large-scale infrastructure

Education

Bachelor’s degree or equivalent practical experience
Master's degree or PhD in Computer Science or related field

Tools

Kubernetes (K8s)
Google Kubernetes Engine (GKE)
Cloud AI platforms

Job description

Google is hiring software engineers to advance AI at scale, delivering LLM inference serving within the Google Distributed Cloud (GDC) Platform. You will work on critical components, from model lifecycle to efficient data loading and dynamic routing, enabling AI across hybrid cloud environments.

You will collaborate with teams across core platform services, tackle complex systems challenges, and contribute to a scalable serving infrastructure that supports innovative AI workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Distributed Cloud AI & LLM
Senior Software Engineer, Distributed Cloud AI & LLM

Google • Town of Montana (WI)

On-site
USD 174,000 - 252,000
Senior Software Engineer, Google Distributed Cloud AI
Senior Software Engineer, Google Distributed Cloud AI

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity
Benefits
Bonus target
GCP AI Architect — Hybrid, LLM & Data Solutions
GCP AI Architect — Hybrid, LLM & Data Solutions

GlobalLogic • United States

Hybrid
USD 150,000 - 160,000
Exciting Projects
Collaborative Environment
Work-Life Balance
+2
Principal Engineer: GKE Platform for AI Inference
Principal Engineer: GKE Platform for AI Inference

Google Inc. • Seattle (WA), Kirkland (WA)

On-site
USD 307,000 - 427,000
Health insurance
Dental insurance
Vision insurance
+5
Hybrid AI Architect & Lead — Edge & Cloud Inference
Hybrid AI Architect & Lead — Edge & Cloud Inference

NIO • San Francisco (CA)

On-site
USD 192,000 - 250,000
Medical plans
Dental & Vision
401(k)
Senior Software Engineer, Google Distributed Cloud AI
Senior Software Engineer, Google Distributed Cloud AI

Google • Town of Montana (WI)

On-site
USD 174,000 - 252,000
ML Inference Engineer - LLM Deployment & Serving
ML Inference Engineer - LLM Deployment & Serving

Google DeepMind • Mountain View (CA)

Hybrid
USD 230,000 - 290,000
Senior ML Systems Engineer – Cloud AI & LLMs
Senior ML Systems Engineer – Cloud AI & LLMs

Google • Seattle (WA)

On-site
USD 262,000 - 364,000
Health Insurance
401(k) with company match
Paid Time Off
+4
GKE AI Inference Platform Architect
GKE AI Inference Platform Architect

Google Inc. • Seattle (WA)

On-site
USD 307,000 - 427,000
Health insurance
401(k) with company match
Paid time off
Senior Software Engineer - Distributed AI & Cloud Infra
Senior Software Engineer - Distributed AI & Cloud Infra

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000