Azure SRE Lead — Cloud Reliability & Automation

SimCorp

Toronto

Hybrid

CAD 113,000 - 142,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health and dental care
Group RRSP/TFSA
Hybrid work policy

Job summary

SimCorp is seeking a Lead Site Reliability Engineer in Toronto to own the reliability, scalability and performance of Azure-hosted environments for our SaaS offering. You will lead onboarding for new and existing client platforms and drive SRE practice across tooling, incident response, and automation.

You will guide automation adoption, mentor engineers, and collaborate with product owners and architects to shape future roadmaps, while participating in on-call weekend support per the needs of

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related field.
  • 5–8+ years in Site Reliability Engineering or Cloud Infrastructure leadership roles.
  • Strong expertise in Microsoft Azure, including production-grade design and operation.
  • Proficiency with IaC tools like Terraform, Bicep, ARM, Ansible.
  • Deep understanding of cloud-native monitoring and incident management frameworks.
  • Hands‑on experience in monitoring and logging tools (Azure Monitor, Application Insights, Log Analytics, Grafana).
  • Familiarity with SimCorp Dimension is a strong plus.
  • Experience managing both onboarding projects and live production operations.
  • Broad knowledge of networking, virtualization, containerization (Kubernetes, Docker).
  • Broad knowledge with Linux and Windows systems, APIs, scripting (PowerShell, Bash), and SQL.
  • Collaborative mindset and ability to work in cross-functional teams.
  • Proven ability to mentor and lead engineers, influence architecture decisions, and manage complexity.
  • Comfort balancing strategic priorities with hands‑on execution.

Responsibilities

  • Own the reliability, scalability, and performance of Azure-hosted environments.
  • Lead operational support and onboarding activities for new and running client platforms.
  • Provide technical leadership in SRE practices, tooling, and incident response.
  • Guide the transition of manual processes into automated, scalable solutions.
  • Lead solutions workshops, engage with vendors, and manage senior stakeholders.
  • Implement and evolve observability frameworks using SLOs, SLIs, and proactive alerting.
  • Oversee disaster recovery planning, incident postmortems, and root cause analysis.
  • Ensure high-quality onboarding delivery through reusable automation pipelines.
  • Mentor and support the growth of junior and senior engineers in your Product Area.
  • Keeping oversight and coordinating routine maintenance, deployments, refreshes, rollbacks, and release Application Upgrades.
  • Execute disaster recovery, configuration management, and infrastructure readiness tasks.
  • Collaborate with product owners and architects to influence platform design and future roadmaps.
  • Provide weekend or on-call support as needed.
  • Contribute to incident, problem, and change management processes.

Skills

Azure expertise
SRE leadership
Terraform
Monitoring tools
Kubernetes
PowerShell
Bash
Linux/Windows

Education

CS degree (Bachelor's or Master's)

Tools

Terraform
Ansible
Grafana
Azure Monitor
Application Insights

Job description

SimCorp is seeking a Lead Site Reliability Engineer in Toronto to own the reliability, scalability and performance of Azure-hosted environments for our SaaS offering. You will lead onboarding for new and existing client platforms and drive SRE practice across tooling, incident response, and automation.

You will guide automation adoption, mentor engineers, and collaborate with product owners and architects to shape future roadmaps, while participating in on-call weekend support per the needs of

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Azure SRE – Hybrid Contract Engineer for Reliability
Azure SRE – Hybrid Contract Engineer for Reliability

Aarorn Technologies Inc • Toronto

Hybrid
CAD 100,000 - 125,000
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Azure SRE Developer
Azure SRE Developer

Aarorn Technologies Inc • Toronto

Hybrid
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Senior SRE – AI-Driven Cloud & Reliability Leader
Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius • Vancouver

Hybrid
CAD 115,000 - 145,000
Senior Site Reliability Engineer – Cloud & Automation Lead
Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc. • Toronto

Remote
CAD 90,000 - 120,000
Digital-first work environment
Collaborative workspaces
Continuous learning opportunities
Lead Site Reliability Engineer
Lead Site Reliability Engineer

SimCorp • Toronto

On-site
CAD 113,000 - 142,000
Health and dental care
Group RRSP/TFSA
Hybrid work policy
Senior SRE Leader: Scale Reliability & Observability
Senior SRE Leader: Scale Reliability & Observability

Rootly • Toronto

On-site
CAD 120,000 - 180,000
Competitive compensation
Comprehensive medical coverage
3 weeks of vacation
+2
Staff Incident Command & Reliability Engineer
Staff Incident Command & Reliability Engineer

IBM • Markham

On-site
CAD 120,000 - 180,000
Senior SRE: Cloud, CI/CD & Incident Leadership (Hybrid)
Senior SRE: Cloud, CI/CD & Incident Leadership (Hybrid)

Morningstar Credit Ratings, LLC • Toronto

On-site
CAD 90,000 - 133,000
Hybrid work environment
Flexible benefits