Site Reliability Engineering Manager

Akkodis

Toronto

Hybrid

CAD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms. The successful candidate will drive SRE best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on‑premises environments.

The role emphasizes technical leadership with people management in a hybrid Toronto environment, and

Qualifications

  • Bachelor’s Degree in Computer Science, Software Engineering, or equivalent experience.
  • Strong understanding of SRE principles, SLIs, SLOs, and error budgets.
  • Hands-on experience with Azure and Kubernetes.
  • Experience with Infrastructure as Code and automation frameworks.
  • Experience with observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics.
  • Strong scripting or programming skills.
  • Excellent communication and leadership abilities.

Responsibilities

  • Lead, mentor, and develop a team of Site Reliability Engineers.
  • Drive reliability, availability, performance, and scalability across critical applications and platforms.
  • Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.
  • Oversee production support operations, on-call processes, incident management, and escalations.
  • Partner with Development, Infrastructure, Security, and DevOps teams to improve service reliability.
  • Drive observability and monitoring initiatives across the organization.
  • Reduce operational effort through automation and self-healing solutions.
  • Lead major incident response activities and support problem management processes.
  • Support capacity planning, resiliency testing, and disaster recovery initiatives.
  • Develop operational standards, runbooks, and knowledge management practices.
  • Recruit, hire, onboard, and develop SRE talent.
  • Promote a culture of continuous improvement, collaboration, and operational excellence.

Skills

SRE
IaC
Automation
Observability
Prod Support

Education

Bachelor's Degree in CS / SW Eng

Tools

Azure
Kubernetes
Terraform

Job description

Manager, Site Reliability Engineering (SRE)

Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.

The successful candidate will drive Site Reliability Engineering best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on-premises environments. This opportunity is ideal for a technical leader with strong experience in cloud operations, production support, automation, and people leadership.

Work Model: Hybrid

About the Opportunity

Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.

What Will the Successful Candidate Do?

The Manager, Site Reliability Engineering (SRE)'s primary responsibilities include, but are not limited to:

  • Lead, mentor, and develop a team of Site Reliability Engineers.
  • Drive reliability, availability, performance, and scalability across critical applications and platforms.
  • Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.
  • Oversee production support operations, on-call processes, incident management, and escalations.
  • Partner with Development, Infrastructure, Security, and DevOps teams to improve service reliability.
  • Drive observability and monitoring initiatives across the organization.
  • Reduce operational effort through automation and self-healing solutions.
  • Lead major incident response activities and support problem management processes.
  • Support capacity planning, resiliency testing, and disaster recovery initiatives.
  • Develop operational standards, runbooks, and knowledge management practices.
  • Recruit, hire, onboard, and develop SRE talent.
  • Promote a culture of continuous improvement, collaboration, and operational excellence.
What the Successful Candidate Needs to Succeed
Must-Have Skills
  • Site Reliability Engineering (SRE)
  • Infrastructure as Code (IaC)
  • Automation & Scripting
  • Observability & Monitoring Tools
  • Production Support Operations
Required Experience
  • 8+ years supporting enterprise applications and distributed systems.
  • 3+ years of experience leading technical teams.
  • Experience implementing Site Reliability Engineering practices.
  • Experience supporting mission-critical production environments.
  • Experience with incident response and problem management.
  • Experience driving automation and operational improvements.
  • Strong stakeholder management and collaboration skills.
Required Qualifications
  • Bachelor’s Degree in Computer Science, Software Engineering, or equivalent experience.
  • Strong understanding of SRE principles, SLIs, SLOs, and error budgets.
  • Hands-on experience with Azure and Kubernetes.
  • Experience with Infrastructure as Code and automation frameworks.
  • Experience with observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics.
  • Strong scripting or programming skills.
  • Excellent communication and leadership abilities.
Preferred Qualifications
  • Experience in fintech, payment processing, or regulated environments.
  • Experience with change management and compliance practices.
  • Knowledge of SDLC and DevOps best practices.
  • Experience with disaster recovery and resiliency testing.
Technical Skills
  • Site Reliability Engineering (SRE) - Must Have
  • Infrastructure as Code (Terraform, ARM, Bicep, etc.)
  • Automation & Scripting
  • DevOps Practices
Soft Skills
  • Strong leadership and coaching abilities
  • Strong problem-solving capabilities
  • Ability to work effectively with cross-functional teams
Additional Requirements
  • Ability to work in a hybrid environment in Toronto.
  • Experience managing production support and operational teams.
  • Ability to lead through major incidents and operational challenges.
Accessibility

At Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone. We foster a workplace where diversity is celebrated and every voice matters.

We encourage applications from individuals of all backgrounds and identities.

#SRE #SiteReliabilityEngineering #Azure #Kubernetes #CloudEngineering #EngineeringManager #DevOps #PlatformEngineering #Observability #Automation #InfrastructureAsCode #TorontoJobs #HybridJobs #TechnologyJobs #HiringNow #Akkodis #CloudOperations #ProductionSupport #LeadershipJobs

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Akkodis • Toronto

Hybrid
CAD 140,000 - 165,000
Bonus
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
Azure SRE Developer
Azure SRE Developer

Aarorn Technologies Inc • Toronto

Hybrid
Site Reliability Engineer (SRE) – Observability
Site Reliability Engineer (SRE) – Observability

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

Hybrid
CAD 75,000 - 95,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

On-site
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
Site Reliability Engineer
Site Reliability Engineer

Kyndryl • Toronto

Hybrid
CAD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Ian Martin Group • Toronto

On-site
CAD 100,000 - 130,000
Senior SRE Engineer (AKS, Azure, Terraform, Kubernetes, and PowerShell.)
Senior SRE Engineer (AKS, Azure, Terraform, Kubernetes, and PowerShell.)

CorGTA • Mississauga

On-site
CAD 88,000 - 94,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

Hybrid
CAD 120,000 - 160,000
DevOps, Kubernetes, and Site Reliability Engineer
DevOps, Kubernetes, and Site Reliability Engineer

Randstad Enterprise • Montreal (administrative region)

On-site
CAD 90,000 - 130,000