Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR)

Toronto

On-site

CAD 155,000 - 165,000

Full time

11 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Quantum Technology Recruiting Inc. (QTR) is seeking a Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance and resiliency of critical apps in Toronto.

The role combines people leadership with hands-on technical delivery across cloud, automation, observability and production reliability. You will drive SRE maturity through SLOs/SLIs, incident response, runbooks, capacity planning and disaster readiness while partnering with

Qualifications

  • 8+ years of experience in senior technical roles supporting complex or distributed systems.
  • 3+ years of people leadership experience.
  • Strong hands-on knowledge of at least one major cloud platform, with Azure preferred; AWS is also acceptable.
  • Strong knowledge of Kubernetes.
  • Near-expert capability with Infrastructure as Code, automation and observability.
  • Strong experience with enterprise observability tools; Dynatrace is preferred, with Datadog, New Relic or AppDynamics also relevant.
  • Linux knowledge.
  • Strong scripting and/or programming capabilities.
  • Strong understanding of SRE principles including SLOs, SLIs, incident management, resiliency and toil reduction.
  • Ability to operate as both a people leader and hands-on technical leader.
  • Excellent communication and stakeholder-management skills.

Responsibilities

  • Lead, coach and develop a team of Site Reliability Engineers.
  • Own the reliability, availability, scalability and performance of assigned business-critical applications and platforms.
  • Help mature SRE practices including SLIs, SLOs, error budgets, incident response, post-incident reviews and continuous reliability improvement.
  • Oversee production operations, incident management, escalation handling and problem management.
  • Drive improvements in observability, monitoring and operational visibility.
  • Partner with Development, DevOps, Infrastructure, Security and Incident Management teams.
  • Reduce operational toil through automation and self-healing capabilities.
  • Lead capacity planning, resiliency testing and disaster recovery readiness.
  • Act as a senior technical escalation point during major production incidents.
  • Establish effective runbooks, documentation standards and knowledge-management practices.
  • Help develop reliability roadmaps for the applications within your portfolio.
  • Support hiring and workforce planning as the SRE organization continues to evolve.

Skills

Senior technical experience
People leadership
Azure or AWS
Kubernetes
IaC & Automation
Observability tools
Linux
Scripting/Programming
SRE principles
Leadership + Hands-on
Communication

Tools

Dynatrace
Datadog
New Relic

Job description

Position: Manager, Site Reliability Engineering (SRE)
Location: Toronto
Salary: $155,000 - $165,000
Posting Type: Open vacancy
Our client is looking for an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and resiliency of business-critical applications and platforms. This is a newly created position reporting to the Director, SRE & DevOps. The successful candidate will combine people leadership with strong hands-on technical expertise, helping mature an evolving SRE function while supporting a large portfolio of critical applications. Our client is looking for someone who can lead and develop an SRE team while remaining technically credible and hands-on across cloud, automation, observability, infrastructure and production reliability.

Description
  • Lead, coach and develop a team of Site Reliability Engineers.
  • Own the reliability, availability, scalability and performance of assigned business-critical applications and platforms.
  • Help mature SRE practices including SLIs, SLOs, error budgets, incident response, post-incident reviews and continuous reliability improvement.
  • Oversee production operations, incident management, escalation handling and problem management.
  • Drive improvements in observability, monitoring and operational visibility.
  • Partner with Development, DevOps, Infrastructure, Security and Incident Management teams.
  • Reduce operational toil through automation and self-healing capabilities.
  • Lead capacity planning, resiliency testing and disaster recovery readiness.
  • Act as a senior technical escalation point during major production incidents.
  • Establish effective runbooks, documentation standards and knowledge-management practices.
  • Help develop reliability roadmaps for the applications within your portfolio.
  • Support hiring and workforce planning as the SRE organization continues to evolve.
Skills
  • 8+ years of experience in senior technical roles supporting complex or distributed systems.
  • 3+ years of people leadership experience.
  • Strong hands-on knowledge of at least one major cloud platform, with Azure preferred; AWS is also acceptable.
  • Strong knowledge of Kubernetes.
  • Near-expert capability with Infrastructure as Code, automation and observability.
  • Strong experience with enterprise observability tools; Dynatrace is preferred, with Datadog, New Relic, AppDynamics or similar also relevant.
  • Linux knowledge – this is an important technical requirement.
  • Strong scripting and/or programming capabilities.
  • Strong understanding of SRE principles including SLOs, SLIs, incident management, resiliency and toil reduction.
  • Ability to operate as both a people leader and hands-on technical leader.
  • Excellent communication and stakeholder-management skills.
Nice to Have
  • Experience within payments, fintech or other highly regulated environments.
  • Exposure to PCI DSS and/or NIST frameworks.
  • Experience working with change-management and compliance processes.
  • Familiarity with SDLC best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Sage Recruiting Inc. • Canada

On-site
CAD 180,000 - 200,000
Unlimited vacation
Comprehensive health and dental benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

Hybrid
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
SRE x 2
SRE x 2

HRB • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Site Reliability Engineer (SRE) – Observability
Site Reliability Engineer (SRE) – Observability

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

Hybrid
CAD 75,000 - 95,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Software Reliability Engineer/Devops
Software Reliability Engineer/Devops

RE Partners • Mississauga

Hybrid
CAD 110,000 - 170,000
Hybrid work model
Senior SRE Engineer — Automation & Platform Reliability
Senior SRE Engineer — Automation & Platform Reliability

RBC • Toronto

On-site
CAD 110,000 - 140,000