Senior Incident Command & Reliability Engineer

IBM

Boston (MA)

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

IBM Software seeks an expert-level Site Reliability Engineer to drive proactive reliability improvements across a multi-cloud platform (AWS, GCP, and Azure). You will own Rootly configurations, incident response, and SLO/SLA frameworks, with 75% focus on engineering and 25% on coaching and post-mortems.

You will work in Cloud Architecture and Reliability - Supportability, collaborating with global teams, setting standards, and communicating complex technical concepts clearly.

Qualifications

  • Bachelor's Degree required.
  • Master's Degree preferred.
  • 10+ years of experience in SRE, incident management, or reliability engineering.
  • Cloud experience with AWS, GCP, or Azure (all three preferred).

Responsibilities

  • Analyze systemic failure patterns and design reliability improvements.
  • Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
  • Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments.
  • Own standards, practices, and incident response across engineering.
  • Edit and review CRCAs for quality and clarity.
  • Develop and deliver training programs; coach teams through post-mortems.
  • Partner with engineering leaders to elevate reliability practices org-wide.

Skills

Rootly configuration
PagerDuty
Jira
Confluence
Slack
Reliability engineering
Incident management
Observability
Kubernetes
CI/CD
Post-mortems

Education

Bachelor's Degree
Master's Degree

Tools

Rootly
PagerDuty
Kafka
AWS

Job description

IBM Software seeks an expert-level Site Reliability Engineer to drive proactive reliability improvements across a multi-cloud platform (AWS, GCP, and Azure). You will own Rootly configurations, incident response, and SLO/SLA frameworks, with 75% focus on engineering and 25% on coaching and post-mortems.

You will work in Cloud Architecture and Reliability - Supportability, collaborating with global teams, setting standards, and communicating complex technical concepts clearly.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Junior Site Reliability Engineer - Automation & Reliability
Junior Site Reliability Engineer - Automation & Reliability

IBM • Tucson (AZ)

On-site
USD 110,000 - 160,000
Site Reliability Engineer: Build Scalable, Resilient Systems
Site Reliability Engineer: Build Scalable, Resilient Systems

QuickCruit • Tucson (AZ), Northern (KY)

Hybrid
USD 120,000 - 150,000
Aspiring Site Reliability Engineer — Cloud & Automation
Aspiring Site Reliability Engineer — Cloud & Automation

IBM • Town of Montana (WI)

On-site
USD 65,000 - 90,000
Healthcare benefits
401(k) plan
Paid time off
Entry Level Site Reliability Engineer - Tucson-AZ
Entry Level Site Reliability Engineer - Tucson-AZ

IBM • Tucson (AZ)

On-site
USD 110,000 - 160,000
Entry Level Site Reliability Engineer IBM · Tucson, AZ
Entry Level Site Reliability Engineer IBM · Tucson, AZ

QuickCruit • Tucson (AZ), Northern (KY)

Hybrid
USD 120,000 - 150,000
Junior Site Reliability Engineer: Build Reliable Systems
Junior Site Reliability Engineer: Build Reliable Systems

IBM Computing • Tucson (AZ)

On-site
USD 100,000 - 150,000
Site Reliability Engineer Intern - Build 24x7 Cloud Reliability
Site Reliability Engineer Intern - Build 24x7 Cloud Reliability

IBM • Durham (NC)

Hybrid
USD 76,000 - 166,000
Site Reliability Engineer Intern: Build Resilient Cloud Systems
Site Reliability Engineer Intern: Build Resilient Cloud Systems

IBM • Austin (TX)

On-site
USD 25,000 - 36,000
SRE Lead: Incident Commander & Reliability Champion
SRE Lead: Incident Commander & Reliability Champion

U.S. Bank • Northern (KY)

Hybrid
USD 112,000 - 131,000
Healthcare
Retirement plan
Paid vacation
+2
ELH Site Reliability Engineer Lowell, SVL, Austin
ELH Site Reliability Engineer Lowell, SVL, Austin

IBM • Lowell (MA)

On-site
USD 85,000 - 110,000