AI Infrastructure Tech Lead — TPU & Kubernetes

Google

Kirkland (WA)

On-site

USD 207,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Life insurance
Disability insurance
401(k) with company match
Paid time off
Sick time

Job summary

Google's software engineers are building the next-generation AI hardware platform and on-premises AI supercomputer. You will define the technical roadmap for inventory and fleet management, emphasizing Temporal orchestration, automated node re-bootstrapping, and self-service remediation.

The role requires deep experience with large-scale distributed systems, Kubernetes and cross-functional collaboration, offering significant impact within Google's cutting-edge infrastructure initiatives.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages (e.g., Python, C, C++, Java, JavaScript).
  • 5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture.
  • 5 years of experience testing, and launching software products.
  • 3 years of experience with software design and architecture.

Responsibilities

  • Define the technical roadmap and architecture for inventory and fleet management, focusing on Temporal orchestration, automated node re-bootstrapping, and self-service remediation.
  • Own architectural and design decisions for the platform, collaborating with leadership and cross-functional teams to prioritize efforts, resolve roadblocks, and manage technical debt.
  • Establish and maintain machine health Service Level Objectives (SLOs), driving continuous improvements in fleet-wide observability and environmental monitoring.
  • Foster a culture of engineering excellence by establishing shared responsibility models with tenant teams and enforcing robust system-level access guardrails.
  • Coach and mentor engineers on the team, guiding their technical development and helping them grow their impact.

Skills

8+ years software development
5 years infra/distributed systems
5 years testing & launching software
3 years software design/architecture

Education

Bachelor's degree or equivalent practical experience

Tools

Kubernetes
ArgoCD
Container orchestration
Temporal workflows

Job description

Google's software engineers are building the next-generation AI hardware platform and on-premises AI supercomputer. You will define the technical roadmap for inventory and fleet management, emphasizing Temporal orchestration, automated node re-bootstrapping, and self-service remediation.

The role requires deep experience with large-scale distributed systems, Kubernetes and cross-functional collaboration, offering significant impact within Google's cutting-edge infrastructure initiatives.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, AI Infra & TPU Systems
Tech Lead, AI Infra & TPU Systems

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Benefits
Tech Lead — On-Prem AI Infrastructure & Orchestration
Tech Lead — On-Prem AI Infrastructure & Orchestration

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Bonus target
Equity
Benefits
Senior AI Infrastructure Field Architect
Senior AI Infrastructure Field Architect

Google • Chicago (IL)

On-site
USD 233,000 - 324,000
Health insurance
401(k) match
Paid time off
Senior AI Infrastructure Field Solutions Architect
Senior AI Infrastructure Field Solutions Architect

Google • Atlanta (GA)

On-site
USD 233,000 - 324,000
Senior AI Infrastructure Architect, Cloud & Data Center
Senior AI Infrastructure Architect, Cloud & Data Center

Google • Sunnyvale (CA)

On-site
USD 233,000 - 324,000
Health Insurance
401(k) with company match
Paid time off 20 days
+4
Tech Lead, TPU AI Infrastructure
Tech Lead, TPU AI Infrastructure

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Bonus target
Equity
Benefits
Technical Lead: Large-Scale AI & Infrastructure
Technical Lead: Large-Scale AI & Infrastructure

Google • Kirkland (WA)

On-site
USD 262,000 - 364,000
Health insurance
Dental insurance
Vision insurance
+4
Senior AI/ML Infra Engineer - TPU Health & Diagnoser
Senior AI/ML Infra Engineer - TPU Health & Diagnoser

Google • Town of Montana (WI)

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+2
Tech Lead, TPU AI Infrastructure
Tech Lead, TPU AI Infrastructure

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Benefits
Group Product Manager, AI Infra Hardware: Lead Next‑Gen AI Systems
Group Product Manager, AI Infra Hardware: Lead Next‑Gen AI Systems

Google • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Health, dental, vision, life, disabled
401(k) with company match
20 days vacation annually
+4