Incident Manager

Aceolution

Pennsylvania

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aceolution is seeking an experienced Incident Manager III to lead critical incident response, problem management, and reliability initiatives across our enterprise environments.

The role requires coordinating cross-functional engineering teams during high-severity production incidents, driving rapid troubleshooting, RCA investigations, and actionable remediation plans, with a focus on observability and automation to improve MTTR and service resilience.

Qualifications

  • Bachelor's degree in Engineering, Computer Science, or a related technical discipline.
  • 3–5 years of experience in SRE, Incident Management, Operations, DevOps, or a similar role.
  • Strong expertise in monitoring tools, logging platforms, and incident response processes.
  • Excellent troubleshooting skills across networking, Linux/Windows servers, cloud infrastructure, and distributed systems.
  • Hands-on experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience with automation using Python, Go, Bash, or PowerShell.
  • Good understanding of microservices, containers, and distributed architectures.
  • Excellent communication, leadership, and stakeholder management skills under pressure.

Responsibilities

  • Lead high-severity P0/P1 incidents as Incident Commander.
  • Coordinate cross-functional engineering teams during critical production incidents.
  • Drive rapid troubleshooting, impact assessment, and resolution decisions.
  • Prepare executive-ready incident summaries and ensure accurate documentation.
  • Apply understanding of application architecture and service dependencies to accelerate mitigation.
  • Lead investigations for recurring and major incidents.
  • Perform Root Cause Analysis with engineering teams.
  • Validate corrective actions and monitor long-term remediation plans.
  • Present findings, lessons learned, and preventive strategies to leadership.
  • Review and approve high-risk production changes.
  • Participate in CAB meetings and validate rollback strategies.
  • Lead maintenance activities and reliability events.
  • Improve runbooks, monitoring strategies, and automation workflows.
  • Contribute ideas around AI to improve efficiency.

Skills

Incident management
SRE
Cloud platforms (AWS/Azure/GCP)
Monitoring & observability
Automation
Communication / leadership

Education

Bachelor's degree in Engineering or Computer Science

Tools

Python
Go
Bash
PowerShell
AWS
Azure
GCP

Job description

We are seeking an experienced Incident Manager III to lead critical incident response, problem management, and operational reliability initiatives. The ideal candidate will have strong technical expertise in cloud technologies, infrastructure, monitoring tools, and incident management, along with the ability to coordinate cross-functional teams during high-severity production incidents.

Key Responsibilities

  • Lead high-severity P0/P1 incidents as the Incident Commander.
  • Coordinate cross-functional engineering teams during critical production incidents.
  • Drive rapid troubleshooting, impact assessment, and resolution decisions.
  • Prepare executive-ready incident summaries and ensure accurate documentation.
  • Apply technical understanding of application architecture and service dependencies to accelerate incident mitigation.
  • Lead investigations for recurring and major incidents.
  • Perform detailed Root Cause Analysis (RCA) with engineering teams.
  • Validate corrective actions and monitor long-term remediation plans.
  • Present findings, lessons learned, and preventive strategies to leadership.

Change Management

  • Review and approve high-risk or complex production changes.
  • Participate in Change Advisory Board (CAB) meetings as required.
  • Validate rollback strategies and operational readiness.
  • Lead change execution for major maintenance activities and reliability events.

Reliability & Operational Excellence

  • Drive initiatives to improve MTTA, MTTR, and overall service reliability.
  • Mentor junior team members on incident response best practices.
  • Partner with SRE and Engineering teams to enhance observability, automation, and resilience.
  • Lead maintenance events, disaster recovery exercises, failover testing, and resilience validation.
  • Review and improve runbooks, monitoring strategies, and automation workflows.
  • Contribute ideas around AI and automation to improve operational efficiency and reduce MTTM/MTTR.

Required Qualifications

  • Bachelor's degree in Engineering, Computer Science, or a related technical discipline.
  • 3–5 years of experience in Site Reliability Engineering (SRE), Incident Management, Operations, DevOps, or a similar role.
  • Strong expertise in monitoring tools, logging platforms, and incident response processes.
  • Excellent troubleshooting skills across networking, Linux/Windows servers, cloud infrastructure, and distributed systems.
  • Hands-on experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Experience with automation using Python, Go, Bash, or PowerShell.
  • Good understanding of microservices, containers, and distributed architectures.
  • Excellent communication, leadership, and stakeholder management skills with the ability to perform under pressure.

Preferred Skills

  • Experience working in enterprise production support environments.
  • Knowledge of ITIL Incident, Problem, and Change Management processes.
  • Exposure to observability platforms and automation frameworks.
  • Experience supporting large-scale cloud-native applications.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Incident And Request Manager
Incident And Request Manager

Compunnel, Inc. • Atlanta (GA)

On-site
USD 90,000 - 120,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Software Engineering Group Manager – Site Reliability Center
Software Engineering Group Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 140,000 - 200,000
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • Reston (VA)

On-site
USD 86,000 - 138,000
Incident Commander: Cloud Reliability & SRE Lead
Incident Commander: Cloud Reliability & SRE Lead

Aceolution • Pennsylvania

On-site
USD 120,000 - 150,000
Systems Operations and Engineering Manager
Systems Operations and Engineering Manager

LOOP • Greenville (SC), Spartanburg (SC), Anderson (SC)

On-site
USD 120,000 - 160,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler-Technologies-29572f8 • Lakewood (CO)

On-site
USD 93,547 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies, Inc. • Plano (TX), Latham (NY), Lubbock (TX), Lakewood (CO)

On-site
USD 93,547 - 150,000
Enterprise SRE – Incident Coordinator
Enterprise SRE – Incident Coordinator

Jobtailor • Missouri

On-site
USD 90,000 - 120,000