Senior Incident Command Engineer – Cloud Reliability

IBM

Ottawa

Hybrid

CAD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

IBM is seeking an experienced Site Reliability Engineer focused on reliability improvements across multi-cloud platforms. You will own incident response, tooling integrations, and continuous learning to reduce outages and improve operational excellence.

The role blends hands-on engineering with strategic program ownership, emphasizing observability, automation, and scalable reliability practices across AWS, GCP, and Azure.

Qualifications

  • 10+ years in SRE/incident management or reliability engineering.
  • Cloud experience with AWS, GCP, or Azure.
  • Experience in large organizations (500+ engineers).
  • Deep expertise with incident management tooling (Rootly, PagerDuty).
  • Strong understanding of distributed systems and failure modes at scale.
  • Kafka/event streaming expertise or quick mastery of complex systems.

Responsibilities

  • Analyze systemic failure patterns to design reliability improvements.
  • Own Rootly configuration, workflows, and integrations with PagerDuty & Jira.
  • Define and maintain SLO/SLA frameworks and use error budgets.
  • Lead incident response standards and process improvements.
  • Edit customer-facing incident documents for clarity.
  • Develop and deliver training and post-mortem coaching.
  • Partner with engineering leaders to elevate reliability practices.

Skills

SRE
Incident management
Reliability engineering
Observability
Kubernetes
CI/CD
Distributed systems
Kafka
Automation
Cloud experience (AWS/GCP/Azure)

Education

Master's Degree

Tools

Rootly
PagerDuty
Confluence
Jira
Slack

Job description

IBM is seeking an experienced Site Reliability Engineer focused on reliability improvements across multi-cloud platforms. You will own incident response, tooling integrations, and continuous learning to reduce outages and improve operational excellence.

The role blends hands-on engineering with strategic program ownership, emphasizing observability, automation, and scalable reliability practices across AWS, GCP, and Azure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Staff Incident Command & Reliability Engineer
Staff Incident Command & Reliability Engineer

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Bedford

On-site
CAD 140,000 - 200,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Toronto

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Edmonton

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Calgary

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Senior SRE – AI-Driven Cloud & Reliability Leader
Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius • Vancouver

Hybrid
CAD 115,000 - 145,000