SRE Engineering Manager - AI Foundry & Global Reliability

Socket.dev

San Jose (CA)

On-site

USD 207,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Google is seeking a Site Reliability Engineer to lead a team responsible for the reliability of large-scale, distributed services. You will own end-to-end availability, build automation to prevent issues, and drive efficiency across systems used in production and critical training pipelines.

You’ll mentor engineers, coordinate after-hours rotations across regions, and shape operational norms while collaborating with SWE, product, and policy teams to uphold high standards and resilient

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8+ years of software development experience in one or more languages.
  • 3+ years designing, analyzing, and troubleshooting distributed systems.
  • 3+ years of experience managing people or teams.
  • 3+ years of experience leading projects.

Responsibilities

  • Manage a team of Software/Systems Engineers and own uptime.
  • Own availability and performance of key services; build automation to prevent recurrence.
  • Mentor the team through credible, quality technical execution.
  • Coordinate on-call rotations across continents with follow-the-sun model.
  • Design and deliver software to improve availability, scalability, latency and efficiency.

Skills

Software development expertise
Distributed systems design
People management
Project leadership

Education

Bachelor’s degree in Computer Science or related field

Job description

Google is seeking a Site Reliability Engineer to lead a team responsible for the reliability of large-scale, distributed services. You will own end-to-end availability, build automation to prevent issues, and drive efficiency across systems used in production and critical training pipelines.

You’ll mentor engineers, coordinate after-hours rotations across regions, and shape operational norms while collaborating with SWE, product, and policy teams to uphold high standards and resilient

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineering Manager, AI Foundry
SRE Engineering Manager, AI Foundry

Google Inc. • San Jose (CA)

On-site
USD 207,000 - 300,000
SRE Engineering Manager: Reliability & Scale Leadership
SRE Engineering Manager: Reliability & Scale Leadership

Google • Raleigh (NC)

On-site
USD 207,000 - 300,000
Equity
Bonus (20% target)
Benefits
SRE Engineering Manager: Scale, Availability & Leadership
SRE Engineering Manager: Scale, Availability & Leadership

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
SRE Engineering Manager - Reliability & Platform Growth
SRE Engineering Manager - Reliability & Platform Growth

Google • Houston (TX)

Hybrid
USD 213,000 - 300,000
SRE Engineering Manager: Lead Reliability & Architecture
SRE Engineering Manager: Lead Reliability & Architecture

Google Inc. • Sunnyvale (CA)

Hybrid
USD 256,000 - 300,000
Equity
Bonus target
Senior SRE Engineering Manager – Lead Uptime
Senior SRE Engineering Manager – Lead Uptime

Google • San Bruno (CA)

On-site
USD 262,000 - 364,000
SRE Engineering Manager — Strategic Growth & Reliability
SRE Engineering Manager — Strategic Growth & Reliability

Epic Games (Portuguese) • Sunnyvale (CA)

On-site
USD 197,000 - 291,000
SRE Manager: AI Infrastructure & Data Intelligence
SRE Manager: AI Infrastructure & Data Intelligence

Socket.dev • San Jose (CA)

On-site
USD 207,000 - 300,000
Staff SRE Engineer — Home IoT & Cloud Reliability
Staff SRE Engineer — Home IoT & Cloud Reliability

Google • San Francisco (CA)

On-site
USD 207,000 - 300,000
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE

Socket.dev • San Jose (CA)

On-site
USD 207,000 - 300,000