Senior SRE: Cloud Reliability & Incident Leader

IBM

Boston (MA)

On-site

USD 95,357 - 177,092

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IBM Software seeks a Staff Site Reliability Engineer to lead Confluent Incident Management & Reliability. You will advance reliability across a multi-cloud platform (AWS, GCP, Azure), driving proactive improvements, building automation, and guiding incident response practices for a global team.

The role blends hands-on engineering with strategic program ownership, including coaching post-mortems, training incident commanders, and standardizing practices across engineering.

Qualifications

  • Bachelor's degree in a relevant field.
  • Master's degree preferred.
  • 10+ years of experience in SRE, incident management, or reliability engineering.
  • Cloud experience with AWS, GCP, or Azure (we run all three).

Responsibilities

  • Analyze systemic failure patterns and design reliability improvements that prevent incident recurrence.
  • Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
  • Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments.
  • Own standards, practices, and continuous improvement of incident response across engineering.
  • Edit and review customer-facing incident documents (CRCAs) to ensure quality and clarity.
  • Develop and deliver training programs; coach teams through post-mortems.
  • Partner with engineering leaders to elevate reliability practices org-wide.
  • Deep experience with observability: metrics, logging, tracing.
  • Kubernetes and container orchestration experience.
  • Understanding of CI/CD pipelines and release processes.
  • Experience driving org-wide process and cultural changes.

Skills

SRE experience
Incident management
Observability
Kubernetes
CI/CD pipelines
Cloud platforms
Multi-cloud

Education

Bachelor's Degree
Master's Degree

Tools

Rootly
PagerDuty
Confluence
Slack
Jira

Job description

IBM Software seeks a Staff Site Reliability Engineer to lead Confluent Incident Management & Reliability. You will advance reliability across a multi-cloud platform (AWS, GCP, Azure), driving proactive improvements, building automation, and guiding incident response practices for a global team.

The role blends hands-on engineering with strategic program ownership, including coaching post-mortems, training incident commanders, and standardizing practices across engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability & Incidents Lead
Senior SRE: Cloud Reliability & Incidents Lead

Illumio • San Jose (CA)

On-site
USD 120,000 - 150,000
Junior Site Reliability Engineer - Automation & Reliability
Junior Site Reliability Engineer - Automation & Reliability

IBM • Tucson (AZ)

On-site
USD 110,000 - 160,000
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Vienna (VA)

On-site
USD 96,000 - 265,000
Medical, dental, vision insurance
401(k) with company match
Paid time off and holidays
+1
SRE Intern: Cloud, Automation & Observability
SRE Intern: Cloud, Automation & Observability

IBM • San Jose (CA)

On-site
USD 30,000 - 47,000
SRE Lead: Reliability, AI-Driven Incident Mastery
SRE Lead: Reliability, AI-Driven Incident Mastery

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 250,000
Lead SRE - Reliability Leader, AI-Driven Incident Response
Lead SRE - Reliability Leader, AI-Driven Incident Response

JPMorganChase • Columbus (OH)

On-site
USD 150,000 - 190,000
SRE Lead: Cloud Reliability, Observability & Automation
SRE Lead: Cloud Reliability, Observability & Automation

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

Hybrid
USD 150,000 - 190,000
Flexible/hybrid work model
Competitive compensation
Health and life insurance
Lead Site Reliability Engineer: AWS Cloud & Automation
Lead Site Reliability Engineer: AWS Cloud & Automation

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Entry Level Site Reliability Engineer - Tucson-AZ
Entry Level Site Reliability Engineer - Tucson-AZ

IBM • Tucson (AZ)

On-site
USD 110,000 - 160,000