Staff Incident Command & Reliability Engineer

IBM

Markham

On-site

CAD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

IBM Canada in Markham is seeking an experienced Site Reliability Engineer to lead reliability initiatives for our multi-cloud platform across AWS, GCP, and Azure. You will drive proactive improvements, own incident response practices, and coach teams through post-mortems.

This role sits within Cloud Architecture and Reliability, offering a follow-the-sun global team and opportunities to shape large-scale systems with strong observability and distributed systems expertise.

Qualifications

  • 10+ years of experience in SRE, incident management, or reliability engineering.
  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Strong understanding of distributed systems and failure modes at scale.
  • Proficiency with incident management tooling and runbooks.

Responsibilities

  • Analyze systemic failure patterns and design reliability improvements.
  • Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
  • Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments.
  • Lead incident response improvements and coordinate post-mortems across teams.
  • Develop and deliver training programs; coach teams through post-mortems.
  • Collaborate with engineering leaders to elevate reliability practices org-wide.

Skills

SRE
Incident management
Reliability engineering
Observability
Post-mortems
Cloud platforms

Education

Master's Degree

Tools

Rootly
PagerDuty
Jira
Confluence
Slack

Job description

IBM Canada in Markham is seeking an experienced Site Reliability Engineer to lead reliability initiatives for our multi-cloud platform across AWS, GCP, and Azure. You will drive proactive improvements, own incident response practices, and coach teams through post-mortems.

This role sits within Cloud Architecture and Reliability, offering a follow-the-sun global team and opportunities to shape large-scale systems with strong observability and distributed systems expertise.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Senior Incident Command Engineer – Cloud Reliability
Senior Incident Command Engineer – Cloud Reliability

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Bedford

On-site
CAD 140,000 - 200,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Calgary

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Toronto

On-site
CAD 140,000 - 190,000
Azure SRE Lead — Cloud Reliability & Automation
Azure SRE Lead — Cloud Reliability & Automation

SimCorp • Toronto

Hybrid
CAD 113,000 - 142,000
Health and dental care
Group RRSP/TFSA
Hybrid work policy
Senior Site Reliability & Infrastructure Lead
Senior Site Reliability & Infrastructure Lead

Aviva Canada • Markham

Hybrid
CAD 125,000 - 175,000
Hybrid flexible work model
Career development opportunities
Wellness programs