Senior Incident Command & Reliability Engineer

IBM

Vancouver

On-site

CAD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

IBM is seeking an experienced Site Reliability Engineer to drive proactive reliability improvements across a multi-cloud platform (AWS, GCP, Azure). You’ll own tooling, post-mortem coaching, and incident response enhancements in a follow-the-sun team within Cloud Architecture and Reliability — Supportability.

Responsibilities include designing SLO/SLA frameworks, partnering with engineering leaders, and advancing observability and release processes.

Qualifications

  • 10+ years of experience in SRE, incident management or reliability engineering.
  • Cloud experience with AWS, GCP, or Azure (at least one).
  • Experience navigating reliability programs at large organizations.
  • Deep expertise with incident tooling such as Rootly or PagerDuty.

Responsibilities

  • Analyze systemic failure patterns and design reliability improvements.
  • Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
  • Define and maintain SLO/SLA frameworks; use error budgets to guide investments.
  • Lead standards and ongoing improvements of incident response across engineering.
  • Develop and deliver training programs; coach teams through post-mortems.
  • Collaborate with engineering leaders to elevate reliability practices org-wide.
  • Work with observability: metrics, logging, tracing; Kubernetes/container orchestration.
  • Understand CI/CD pipelines and release processes; communicate clearly in writing.

Skills

SRE experience
Incident management
Reliability engineering
Cloud platforms (AWS/GCP/Azure)
Observability
Kubernetes

Education

Master's Degree

Tools

Rootly
PagerDuty
Confluence
Slack
Jira
CI/CD tooling

Job description

IBM is seeking an experienced Site Reliability Engineer to drive proactive reliability improvements across a multi-cloud platform (AWS, GCP, Azure). You’ll own tooling, post-mortem coaching, and incident response enhancements in a follow-the-sun team within Cloud Architecture and Reliability — Supportability.

Responsibilities include designing SLO/SLA frameworks, partnering with engineering leaders, and advancing observability and release processes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Incident Command Engineer – Cloud Reliability
Senior Incident Command Engineer – Cloud Reliability

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Staff Incident Command & Reliability Engineer
Staff Incident Command & Reliability Engineer

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Bedford

On-site
CAD 140,000 - 200,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Calgary

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Edmonton

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Toronto

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Senior SRE – AI-Driven Cloud & Reliability Leader
Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius • Vancouver

Hybrid
CAD 115,000 - 145,000