Senior Manager, Site Reliability Engineering (SRE)

Dawninfotek

Toronto

On-site

CAD 150,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Dawninfotek is seeking a hands-on Senior Manager to lead Site Reliability Engineering (SRE), Service Delivery, and Infrastructure patches for a Digital Banking platform. You will drive reliability, scale, and incident response across AWS, OpenShift, and legacy systems, guiding a team of 8–10 engineers while partnering with security and product teams.

You'll champion observability, CI/CD improvements, postmortems, and zero-downtime patching to ensure secure, always-on banking services for

Qualifications

  • Hands-on leadership of SRE teams in large-scale environments.
  • Experience driving incident management and RCA programs.
  • Strong collaboration with security, platform, and product teams.

Responsibilities

  • Act as senior escalation point for on-call teams and incidents.
  • Lead major incident response and RCA efforts.
  • Define reliability, scalability, and availability practices for large digital banking platforms.
  • Oversee patching and maintenance across cloud and on‑prem environments.
  • Champion observability, monitoring, and alerting to reduce customer impact.
  • Mentor 8–10 SREs and drive continuous improvement.

Skills

SRE leadership
Incident management
Observability & monitoring
CI/CD & deployment
Cloud platforms (AWS)
OpenShift
Linux
Stakeholder communication

Education

Bachelor's degree in a technical field

Tools

Dynatrace
OpenSearch
Prometheus
Grafana
OpenShift
AWS
Linux
WebSphere

Job description

Senior Manager, SiteReliabilityEngineering(SRE)
Contract to hire for a BANK
Role Overview

We are seeking a hands-on and strategic Senior Manager to lead our Site Reliability Engineering (SRE), Service Delivery, and Infrastructure Patching teams supporting the Digital Banking Platform. This role is crucial to our mission of providing always-on, secure, and high-performing banking services for millions of customers.

KeyResponsibilities
TechnicalLeadership&IncidentManagement
  • Act as the senior technical escalation point for on‑call teams, diagnosing and resolving complex infrastructure, cloud, and application issues.
  • Lead major incident response efforts, ensuring rapid restoration and comprehensive root cause analysis (RCA).
  • Collaborate across engineering, platform, and security to troubleshoot issues spanning full‑stack environments (cloud, container, and legacy platforms).
  • Maintain high availability and performance of digital banking applications (primarily AWS, OpenShift, Linux, with some legacy WebSphere).
  • Champion proactive monitoring, observability, and alerting (e.g., Dynatrace, OpenSearch, Prometheus, Grafana).
SRE&ReliabilityEngineering
  • Define and implement best practices for reliability, scalability, and availability tailored to large‑scale digital banking.
  • Continuously improve CI/CD pipelines, release automation, and deployment practices.
  • Drive rigorous postmortem analysis and a culture of blameless continuous improvement.
  • Optimize for scalability, redundancy, and resilience—minimizing customer impact from incidents.
Infrastructure&Patching
  • Oversee patching and maintenance for cloud and on‑prem environments (AWS, OpenShift, Red Hat VMs, some WebSphere).
  • Ensure zero‑downtime patching strategies and automation to mitigate operational risk and security vulnerabilities.
  • Partner with security teams to enforce compliance, harden platforms, and remediate vulnerabilities.
TeamLeadership&ProcessImprovement
  • Lead, mentor, and grow a high‑performing team of 8–10 SREs and service engineers.
  • Drive a culture of ownership, operational excellence, and continuous learning.
  • Establish and enforce best practices for incident management, operational documentation, and process automation.
  • Collaborate with development, infrastructure, and product teams to enhance observability, deployment, and proactive issue detection
RequiredSkills
  • Exceptional hands‑on troubleshooting skills in complex, distributed, or high‑availability technical environments.
  • Experience in observability, monitoring, and incident management for critical platforms.
  • Demonstrated leadership in technical settings—may include leading projects, initiatives, or mentoring teams, even if not previously a formal people manager.
  • Excellent communicator, able to translate technical detail for both engineers and executives

Bachelor’s degree in a technical field

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Info Way Solutions • Vancouver

On-site
CAD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer (SRE) – Automation & Observability
Senior Site Reliability Engineer (SRE) – Automation & Observability

Tech Talent International • Montreal (administrative region)

Hybrid
CAD 110,000 - 120,000
9% bonus
3–5 weeks vacation
RRSP contribution
+2
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Site Reliability Engineer (Linux / Cloud Infrastructure)
Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT Group • Montreal

On-site
CAD 80,000 - 100,000
Sr Incident & Reliability Manager
Sr Incident & Reliability Manager

Paymentus • Richmond Hill

On-site
CAD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

Hybrid
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4