Lead Site Reliability Engineering to support and enhance Azure-based developer platforms and cloud-native applications

S I Systems

Toronto

On-site

CAD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

S I Systems seeks a Lead Site Reliability Engineer to enhance Azure-based platforms and cloud-native apps in a financial services context. You will blend SRE and platform engineering to improve reliability, observability, and deploy automation across DEV/UAT/PROD.

Lead incident response, tuning metrics like SLOs/SLIs, and drive CI/CD improvements with GitHub Actions. On-site in Toronto, 4 days per week, with a focus on Azure services and production readiness.

Qualifications

  • 5+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support
  • Hands-on experience with Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure Functions
  • Production support experience with incident response, troubleshooting, problem management, and operational support processes
  • Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms
  • Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline

Responsibilities

  • Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments
  • Respond to incidents and lead triage activities to restore service and resolve issues
  • Support deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflows
  • Configure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud services
  • Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log Analytics
  • Support onboarding of new applications and teams to the developer platform
  • Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation

Skills

SRE
Platform engineering
Cloud operations
DevOps
Production support
Microsoft Azure
CI/CD pipelines
GitHub Actions

Education

Post-secondary education in Computer Science/Software Engineering/IT

Tools

Azure Container Apps
Azure Active Directory (Entra ID)
Key Vault
Azure SQL
API Management (APIM)
Azure Functions
Datadog
Azure Monitor
Log Analytics

Job description

Lead Site Reliability Engineering to support and enhance Azure-based developer platforms and cloud-native applications

Our financial services client is seeking a Lead, Site Reliability Engineering (5+ years) to support and enhance Azure-based developer platforms and cloud-native applications

Join a team focused on improving the reliability, availability, and operability of enterprise developer platforms and internal applications. This role combines Site Reliability Engineering, platform engineering, and production support across Azure-based environments, CI/CD pipelines, observability tooling, and cloud-native services. The position supports critical application environments while contributing to automation, incident response, operational readiness, and continuous reliability improvement initiatives.

Must Haves
  • 5+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support
  • Hands-on experience with Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure Functions
  • Production support experience with incident response, troubleshooting, problem management, and operational support processes
  • Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms
  • Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline
Nice to Have
  • Experience with container apps and Kubernetes container orchestration platforms
  • Experience supporting enterprise developer platforms or internal platform engineering teams
  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, and error budgets
  • Exposure to AI, LLM, or agent-based technology platforms
  • Azure, Network, DevOps, or cloud-related certifications
Responsibilities
  • Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments
  • Respond to incidents and lead triage activities to restore service and resolve issues
  • Support deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflows
  • Configure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud services
  • Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log Analytics
  • Support onboarding of new applications and teams to the developer platform
  • Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation

Our financial services client is seeking a Lead, Site Reliability Engineering (5+ years) to support and enhance Azure-based developer platforms and cloud-native applications

Join a team focused on improving the reliability, availability, and operability of enterprise developer platforms and internal applications. This role combines Site Reliability Engineering, platform engineering, and production support across Azure-based environments, CI/CD pipelines, observability tooling, and cloud-native services. The position supports critical application environments while contributing to automation, incident response, operational readiness, and continuous reliability improvement initiatives.

Permanent, Toronto, 4 days/ week on site

Must Haves
  • 5+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support
  • Hands-on experience with Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure Functions
  • Production support experience with incident response, troubleshooting, problem management, and operational support processes
  • Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms
  • Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline
Nice to Have
  • Experience with container apps and Kubernetes container orchestration platforms
  • Experience supporting enterprise developer platforms or internal platform engineering teams
  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, and error budgets
  • Exposure to AI, LLM, or agent-based technology platforms
  • Azure, Network, DevOps, or cloud-related certifications
Responsibilities
  • Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments
  • Respond to incidents and lead triage activities to restore service and resolve issues
  • Support deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflows
  • Configure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud services
  • Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log Analytics
  • Support onboarding of new applications and teams to the developer platform
  • Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Johnson Controls, Inc. • Richmond Hill

On-site
CAD 120,000 - 180,000
Lead Azure SRE for Cloud-Native Platforms
Lead Azure SRE for Cloud-Native Platforms

S I Systems • Toronto

On-site
CAD 120,000 - 180,000
DevOps Support
DevOps Support

Linux Association of Canada • Toronto

Hybrid
CAD 90,000 - 130,000
Senior Azure DevOps Engineer
Senior Azure DevOps Engineer

High Tech Genesis Inc. • Canada

Hybrid
CAD 120,000 - 180,000
Software Engineer (ID#5509)
Software Engineer (ID#5509)

New Value Solutions • Vancouver

On-site
CAD 110,000 - 150,000
Lead DevOps Engineer
Lead DevOps Engineer

Veriday Inc • Toronto

Hybrid
Senior Lead Software Engineer - Squad Tech Lead to lead Java and Spring Boot microservices and high-scale API development - 1757
Senior Lead Software Engineer - Squad Tech Lead to lead Java and Spring Boot microservices and high-scale API development - 1757

S I Systems • Toronto

Hybrid
CAD 140,000 - 190,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Akkodis • Toronto

Hybrid
CAD 140,000 - 165,000
Bonus
Benefits
Lead Cloud Operations Engineer – Azure & Kubernetes
Lead Cloud Operations Engineer – Azure & Kubernetes

Myticas Consulting • Toronto

On-site
CAD 140,000 - 170,000