AI Infrastructure SRE Manager – Scale & Reliability

Google

San Jose (CA)

On-site

USD 207,000 - 300,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Google's Core AI Foundations SRE team is seeking an experienced Site Reliability Engineer to lead a North American SRE group, partnering with development teams and our Sydney counterpart to ensure exceptional reliability for Google's frontier AI capabilities.

The role combines software and systems engineering to design, build, and operate large-scale, fault-tolerant infrastructure, with opportunities to mentor engineers, drive roadmaps, and advance AI-driven automation across products such as

Qualifications

  • Bachelor’s degree in Computer Science, related field, or equivalent practical experience.
  • 8 years of software development experience in one or more programming languages (C++, Java, Python) or with data structures/algorithms.
  • 3 years designing, analyzing, and troubleshooting large-scale distributed systems.
  • 3 years managing and growing engineering teams, including performance management and career development.

Responsibilities

  • Hire, develop, and mentor a high-performing SRE team to support career growth and team health.
  • Set team goals, prioritize resources, and define technical roadmaps aligned with partner teams and key stakeholders.
  • Guide the architecture and review of resilient, high-performance systems powering core AI infrastructure.
  • Drive incident response, maintain high reliability standards, and actively automate operational toil.
  • Partner across development teams to align technical direction while leveraging AI to accelerate team productivity.

Skills

C++
Java
Python
Distributed systems
People management

Education

Bachelor’s degree in CS or related field

Job description

Google's Core AI Foundations SRE team is seeking an experienced Site Reliability Engineer to lead a North American SRE group, partnering with development teams and our Sydney counterpart to ensure exceptional reliability for Google's frontier AI capabilities.

The role combines software and systems engineering to design, build, and operate large-scale, fault-tolerant infrastructure, with opportunities to mentor engineers, drive roadmaps, and advance AI-driven automation across products such as

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineering Manager II, Site Reliability Engineering, Data Intelligence
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Senior SRE & Data Intelligence Engineering Manager
Senior SRE & Data Intelligence Engineering Manager

Google Inc. • San Jose (CA), Northern (KY)

Hybrid
USD 207,000 - 300,000
Equity
Benefits
Bonus target
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence

Google Inc. • San Jose (CA), Northern (KY)

Hybrid
USD 207,000 - 300,000
Equity
Benefits
Bonus target
SRE Engineering Manager, AI Foundry
SRE Engineering Manager, AI Foundry

Google Inc. • San Jose (CA)

On-site
USD 207,000 - 300,000
Senior Staff SRE: Scale, Reliability & Automation
Senior Staff SRE: Scale, Reliability & Automation

Google • New York (NY)

On-site
USD 262,000 - 364,000
Director of Cloud SRE and Global Infrastructure
Director of Cloud SRE and Global Infrastructure

Google • Sunnyvale (CA)

On-site
USD 307,000 - 427,000
Senior SRE Engineer: Scale, Reliability & Automation
Senior SRE Engineer: Scale, Reliability & Automation

engineeringjobs.net, Inc. • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Senior SRE Engineer: Scale, Reliability & Automation
Senior SRE Engineer: Scale, Reliability & Automation

Google • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Senior SRE: AI-Driven, Resilient Data Pipelines
Senior SRE: AI-Driven, Resilient Data Pipelines

Google • Pittsburgh

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits