Site Reliability Manager, Data Center Networking, SRE

Google Canada

Southwestern Ontario

On-site

CAD 216,000 - 221,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Google's Site Reliability Engineering (SRE) team in Waterloo, Ontario is seeking a Site Reliability Manager, Data Center Networking to lead a high-performing team at the intersection of software engineering and large-scale infrastructure.

As the Site Reliability Manager for Data Center Networking, you will drive incident detection and mitigation improvements (MTTD/MTTM), mentor Tech Leads, and shape network product design to ensure reliability across Google Cloud.

Qualifications

  • Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development, including data structures and algorithms.
  • 3 years of experience managing people or teams, leading projects, and troubleshooting distributed systems.
  • Master's degree or PhD in Computer Science or Engineering, or a related field (preferred).

Responsibilities

  • Build and sustain a mission-first culture across multiple locations by scaling leadership through trusted Tech Leads and domain experts
  • Actively prioritize the team's workload to ensure sustained high performance and healthy on-call rotations
  • Own execution of the team's strategic efforts, with a focus on drastically improving MTTD and MTTM for incidents through advanced signaling and tooling
  • Integrate new signals into auto-mitigation systems to reduce incident impact and manual intervention
  • Set the bar for developer excellence across the SDN ecosystem, influencing the design and safe rollout of new network products (NPIs)
  • Partner with PLANET and sibling SRE shards to define and monitor Network SLOs and establish end-to-end repair coverage for network infrastructure
  • Co-own blameless postmortems to drive systemic improvements following incidents

Skills

Software engineering
Distributed systems
Leadership

Education

Bachelor's degree in Computer Science or related field
Master's degree or PhD in CS or Engineering (preferred)

Job description

Google's Site Reliability Engineering (SRE) team in Waterloo, Ontario is looking for a Site Reliability Manager, Data Center Networking to lead a high-performing team at the intersection of software engineering and large-scale infrastructure. This is a senior leadership role within Google's broader SRE organization, where your work directly shapes the reliability and performance of systems that serve millions of GCP customers worldwide.

In this position, you'll be responsible for building a mission-first culture, scaling your leadership through trusted Tech Leads and domain experts, and driving meaningful improvements to incident detection and mitigation across Software-Defined Networking (SDN) infrastructure. It's a role that demands both deep technical expertise and the people leadership skills to move a complex, distributed organization forward.

About the Role: Site Reliability Manager, Data Center Networking

As the Site Reliability Manager for Data Center Networking, you'll serve as the ultimate execution owner for your team's strategic initiatives. Your primary focus will be on dramatically improving Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) for incidents — leveraging advanced signaling, tooling, and the integration of new signals into auto-mitigation systems. You'll also set the standard for developer excellence across the SDN ecosystem, influencing the design and rollout of new network products to ensure they're introduced safely and deliver high reliability.

Collaboration is central to this role. You'll partner closely with the PLANET team and sibling SRE shards to define and monitor Network Service Level Objectives (SLOs), co-own blameless postmortems, and establish end-to-end repair coverage for network infrastructure. Google's SRE culture is grounded in intellectual curiosity, psychological safety, and a commitment to eliminating toil through automation and systems thinking.

Benefits and Salary

This role offers a salary range of CAD $216,000 – $221,000 per year, plus a 20% bonus target, equity, and a comprehensive benefits package. Individual compensation is determined by job-related skills, experience, and relevant education or training. For full details on Google's benefits, visit their official careers page.

Job Details

Company: Google

Location: Waterloo, ON, Canada

Requisition ID: 92399333910946502

Pay: CAD $216,000 – $221,000 per year + 20% bonus target + equity + benefits

Responsibilities

This role spans strategic leadership, technical influence, and operational ownership. You'll be expected to drive reliability improvements at scale while mentoring and empowering the people around you — all within a culture that values blameless learning and continuous improvement.

  • Build and sustain a cohesive, mission-first culture across multiple locations by scaling leadership through trusted Tech Leads and domain experts
  • Actively prioritize the team's workload to ensure sustained high performance and healthy on-call rotations
  • Own execution of the team's strategic efforts, with a focus on drastically improving MTTD and MTTM for incidents through advanced signaling and tooling
  • Integrate new signals into auto-mitigation systems to reduce incident impact and manual intervention
  • Set the bar for developer excellence across the SDN ecosystem, influencing the design and safe rollout of new network products (NPIs)
  • Partner with PLANET and sibling SRE shards to define and monitor Network SLOs and establish end-to-end repair coverage for network infrastructure
  • Co-own blameless postmortems to drive systemic improvements following incidents
Requirements / Skills

Google is looking for a leader who brings both deep software engineering expertise and proven experience managing technical teams in complex, distributed environments. The ideal candidate thrives in ambiguity, leads with curiosity, and has a track record of improving system reliability at scale.

  • Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience
  • 8 years of experience in software development, including work with data structures and algorithms
  • 3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems
  • Master's degree or PhD in Computer Science, Engineering, or a related field is preferred

Bachelor's degree in Computer Science or a related technical field or equivalent practical experience. 8 years of experience in software development, and with data structures and algorithms. 3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems. Master's degree or PhD in Computer Science or Engineering, or a related field (preferred).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Manager, Data Center Networking, SRE
Site Reliability Manager, Data Center Networking, SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Bonus target
Benefits
Site Reliability Manager, Data Center Networking, SRE
Site Reliability Manager, Data Center Networking, SRE

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Benefits
Bonus target 20%
Staff Software Developer, Protected Data SRE
Staff Software Developer, Protected Data SRE

Google Canada • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Benefits package
Bonus target
Staff Systems Developer Manager, AlphaNet Core SRE
Staff Systems Developer Manager, AlphaNet Core SRE

Socket.dev • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Staff Systems Developer Manager, AlphaNet Core SRE
Staff Systems Developer Manager, AlphaNet Core SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Staff Systems Developer Manager, AlphaNet Core SRE
Staff Systems Developer Manager, AlphaNet Core SRE

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Benefits package
Bonus target 20%
Senior Systems & Reliability Manager (SRE)
Senior Systems & Reliability Manager (SRE)

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Benefits package
Bonus target 20%
Cloud Platform and Infrastructure Specialist, Google Cloud
Cloud Platform and Infrastructure Specialist, Google Cloud

Google Canada • Toronto

On-site
CAD 170,000 - 174,000
Cloud Technical Solutions Developer, Storage and Databases
Cloud Technical Solutions Developer, Storage and Databases

Google Canada • Southwestern Ontario

Hybrid
CAD 144,000 - 147,000
Equity
Benefits package
Site Reliability Engineer
Site Reliability Engineer

Pacer Group • Toronto

Hybrid
CAD 69,000 - 76,000