Incident Commander: Cloud Reliability & SRE Lead

Aceolution

Pennsylvania

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aceolution is seeking an experienced Incident Manager III to lead critical incident response, problem management, and reliability initiatives across our enterprise environments.

The role requires coordinating cross-functional engineering teams during high-severity production incidents, driving rapid troubleshooting, RCA investigations, and actionable remediation plans, with a focus on observability and automation to improve MTTR and service resilience.

Qualifications

  • Bachelor's degree in Engineering, Computer Science, or a related technical discipline.
  • 3–5 years of experience in SRE, Incident Management, Operations, DevOps, or a similar role.
  • Strong expertise in monitoring tools, logging platforms, and incident response processes.
  • Excellent troubleshooting skills across networking, Linux/Windows servers, cloud infrastructure, and distributed systems.
  • Hands-on experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience with automation using Python, Go, Bash, or PowerShell.
  • Good understanding of microservices, containers, and distributed architectures.
  • Excellent communication, leadership, and stakeholder management skills under pressure.

Responsibilities

  • Lead high-severity P0/P1 incidents as Incident Commander.
  • Coordinate cross-functional engineering teams during critical production incidents.
  • Drive rapid troubleshooting, impact assessment, and resolution decisions.
  • Prepare executive-ready incident summaries and ensure accurate documentation.
  • Apply understanding of application architecture and service dependencies to accelerate mitigation.
  • Lead investigations for recurring and major incidents.
  • Perform Root Cause Analysis with engineering teams.
  • Validate corrective actions and monitor long-term remediation plans.
  • Present findings, lessons learned, and preventive strategies to leadership.
  • Review and approve high-risk production changes.
  • Participate in CAB meetings and validate rollback strategies.
  • Lead maintenance activities and reliability events.
  • Improve runbooks, monitoring strategies, and automation workflows.
  • Contribute ideas around AI to improve efficiency.

Skills

Incident management
SRE
Cloud platforms (AWS/Azure/GCP)
Monitoring & observability
Automation
Communication / leadership

Education

Bachelor's degree in Engineering or Computer Science

Tools

Python
Go
Bash
PowerShell
AWS
Azure
GCP

Job description

Aceolution is seeking an experienced Incident Manager III to lead critical incident response, problem management, and reliability initiatives across our enterprise environments.

The role requires coordinating cross-functional engineering teams during high-severity production incidents, driving rapid troubleshooting, RCA investigations, and actionable remediation plans, with a focus on observability and automation to improve MTTR and service resilience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Incident Manager
Incident Manager

Aceolution • Pennsylvania

On-site
USD 120,000 - 150,000
Remote Incident Commander & Reliability Lead
Remote Incident Commander & Reliability Lead

Cacheflow • United States

Remote
USD 103,000 - 146,000
Comprehensive benefits
Equity opportunities
Performance bonus eligibility
Cloud Incident & Reliability Lead (DevSecOps)
Cloud Incident & Reliability Lead (DevSecOps)

Peraton • Reston (VA)

On-site
USD 86,000 - 138,000
Senior SRE: Cloud Reliability & Incidents Lead
Senior SRE: Cloud Reliability & Incidents Lead

Illumio • San Jose (CA)

On-site
USD 120,000 - 150,000
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Vienna (VA)

On-site
USD 96,000 - 265,000
Medical, dental, vision insurance
401(k) with company match
Paid time off and holidays
+1
Cloud & Enterprise Incident Command Lead
Cloud & Enterprise Incident Command Lead

Peraton • United States

On-site
USD 86,000 - 138,000
Medical and dental insurance
401(k) plan
Paid time off (PTO)
Senior Incident Commander & Reliability Lead
Senior Incident Commander & Reliability Lead

Databricks Inc. • United States

On-site
USD 103,000 - 146,000
Senior Incident Manager, 24x7 Cloud Reliability | Hybrid
Senior Incident Manager, 24x7 Cloud Reliability | Hybrid

Optomi • Baltimore (MD)

Hybrid
USD 85,000 - 120,000
Remote Incident & Problem Manager - Tier II
Remote Incident & Problem Manager - Tier II

F3 Design • United States

Remote
USD 93,000 - 109,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000