Lead Site Reliability Engineering to support and enhance Azure-based developer platforms and cloud-native applications
Our financial services client is seeking a Lead, Site Reliability Engineering (5+ years) to support and enhance Azure-based developer platforms and cloud-native applications
Join a team focused on improving the reliability, availability, and operability of enterprise developer platforms and internal applications. This role combines Site Reliability Engineering, platform engineering, and production support across Azure-based environments, CI/CD pipelines, observability tooling, and cloud-native services. The position supports critical application environments while contributing to automation, incident response, operational readiness, and continuous reliability improvement initiatives.
Must Haves
- 5+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support
- Hands-on experience with Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure Functions
- Production support experience with incident response, troubleshooting, problem management, and operational support processes
- Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms
- Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline
Nice to Have
- Experience with container apps and Kubernetes container orchestration platforms
- Experience supporting enterprise developer platforms or internal platform engineering teams
- Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, and error budgets
- Exposure to AI, LLM, or agent-based technology platforms
- Azure, Network, DevOps, or cloud-related certifications
Responsibilities
- Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments
- Respond to incidents and lead triage activities to restore service and resolve issues
- Support deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflows
- Configure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud services
- Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log Analytics
- Support onboarding of new applications and teams to the developer platform
- Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation
Our financial services client is seeking a Lead, Site Reliability Engineering (5+ years) to support and enhance Azure-based developer platforms and cloud-native applications
Join a team focused on improving the reliability, availability, and operability of enterprise developer platforms and internal applications. This role combines Site Reliability Engineering, platform engineering, and production support across Azure-based environments, CI/CD pipelines, observability tooling, and cloud-native services. The position supports critical application environments while contributing to automation, incident response, operational readiness, and continuous reliability improvement initiatives.
Permanent, Toronto, 4 days/ week on site
Must Haves
- 5+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support
- Hands-on experience with Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure Functions
- Production support experience with incident response, troubleshooting, problem management, and operational support processes
- Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms
- Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline
Nice to Have
- Experience with container apps and Kubernetes container orchestration platforms
- Experience supporting enterprise developer platforms or internal platform engineering teams
- Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, and error budgets
- Exposure to AI, LLM, or agent-based technology platforms
- Azure, Network, DevOps, or cloud-related certifications
Responsibilities
- Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments
- Respond to incidents and lead triage activities to restore service and resolve issues
- Support deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflows
- Configure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud services
- Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log Analytics
- Support onboarding of new applications and teams to the developer platform
- Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation