Senior Incident Manager

System One

Knoxville (TN)

On-site

USD 110,000 - 160,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

System One seeks a Senior Incident Manager to lead the response and resolution of high impact technology incidents in a 24x7 enterprise command center. You will guide cross-functional teams across AWS, networks, and Unix/Linux, ensuring timely resolution and clear stakeholder updates.

The role emphasizes leadership in incident bridges, risk assessment, and post-incident reviews, rather than implementing fixes directly, with close executive communication during critical events.

Qualifications

  • 5+ years of Major Incident Management / Enterprise Incident Management experience.
  • Experience leading high severity incident bridges or enterprise command center calls.
  • Strong understanding of enterprise networking and infrastructure; AWS knowledge; Unix/Linux expertise.
  • Excellent executive level verbal and written communication skills.

Responsibilities

  • Lead major enterprise incidents from initial identification through mitigation, recovery, and closure.
  • Manage business impacting incidents involving applications, AWS/cloud services, networks, and enterprise infrastructure.
  • Lead incident bridges that may include 50–100 technical participants, stakeholders, and senior leaders.
  • Perform technical and business impact triage and establish clear incident priorities.
  • Coordinate application, infrastructure, network, cloud, Unix/Linux, and other SME teams during incident response.
  • Use AWS and non AWS monitoring tools to gather evidence and support technical investigation.
  • Trace transactions, dependencies, and system to system communication across enterprise environments.
  • Assess technical risk, business impact, mitigation options, and recovery status.
  • Recommend mitigation and remediation actions to the responsible engineering teams.
  • Ensure incidents progress within applicable enterprise SLAs.
  • Maintain clear, concise communication with technical teams, business stakeholders, and senior leadership.
  • Support root cause analysis, lessons learned, and continuous improvement activities following major incidents.
  • Maintain focus, accountability, and effective decision making throughout high pressure incidents.

Skills

Major Incident Management
Incident Bridges leadership
Executive communication
Cross-functional coordination
Unix/Linux knowledge

Education

Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field

Tools

AWS monitoring tools

Job description

Job Title: Senior Incident Manager

Job Type: Permanent Full Time

Overview

System One is seeking a Senior Incident Manager to lead the response and resolution of high impact technology incidents within a 24x7x365 enterprise command center environment. This role is responsible for managing incidents from both technical and business perspectives across AWS, enterprise applications, networks, infrastructure, and Unix/Linux environments.

The Incident Manager will lead large incident bridges, coordinate technical investigation across multiple teams, assess business impact and risk, evaluate mitigation and recovery options, and ensure incidents progress toward resolution within established SLAs. The role requires enough technical depth to ask the right questions, trace dependencies and transaction flows, interpret monitoring information, and recommend appropriate actions while keeping large response teams focused during critical events.

This is an incident leadership and technical advisory role, rather than a position responsible for directly implementing infrastructure fixes.

Responsibilities
  • Lead major enterprise incidents from initial identification through mitigation, recovery, and closure.
  • Manage business impacting incidents involving applications, AWS/cloud services, networks, and enterprise infrastructure.
  • Lead incident bridges that may include 50–100 technical participants, stakeholders, and senior leaders.
  • Perform technical and business impact triage and establish clear incident priorities.
  • Coordinate application, infrastructure, network, cloud, Unix/Linux, and other SME teams during incident response.
  • Use AWS and non AWS monitoring tools to gather evidence and support technical investigation.
  • Trace transactions, dependencies, and system to system communication across enterprise environments.
  • Assess technical risk, business impact, mitigation options, and recovery status.
  • Recommend mitigation and remediation actions to the responsible engineering teams.
  • Ensure incidents progress within applicable enterprise SLAs.
  • Maintain clear, concise communication with technical teams, business stakeholders, and senior leadership.
  • Support root cause analysis, lessons learned, and continuous improvement activities following major incidents.
  • Maintain focus, accountability, and effective decision making throughout high pressure incidents.
Requirements
  • 5+ years of Major Incident Management / Enterprise Incident Management experience.
  • Proven experience leading high severity incident bridges or enterprise command center calls.
  • Ability to coordinate large, cross functional technical teams during high impact production incidents.
  • Strong understanding of enterprise networking and infrastructure, including routers, switches, BGP, TCP/IP, and DNS.
  • Working knowledge of AWS environments and cloud related production troubleshooting.
  • Strong Unix/Linux knowledge and broad enterprise infrastructure awareness.
  • Ability to understand application integrations, dependencies, transaction flows, and system to system communication.
  • Strong technical triage skills with the ability to gather evidence and identify the right SME or technical team.
  • Experience assessing business impact, technical risk, mitigation options, and recovery status.
  • Root cause analysis and post incident review experience.
  • Excellent executive level verbal and written communication skills.
  • Strong facilitation, leadership, composure, and decision making under pressure.
  • Ability to drive accountability and maintain focus across large incident response teams.
  • Experience operating within a 24x7 enterprise production environment.
Desired Skillset:
  • AWS Solutions Architect Associate certification and experience supporting large, regulated enterprise environments.
  • Experience communicating directly with executives during critical incidents is also highly valuable.
Educational Requirements:

Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field.

Ref: #404-IT Pittsburgh

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Incident Leader – 24/7 Enterprise Response
Senior Incident Leader – 24/7 Enterprise Response

System One • Knoxville (TN)

On-site
USD 110,000 - 160,000
Incident Manager
Incident Manager

Insight Global • Jersey City (NJ)

On-site
USD 110,000 - 140,000
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • Reston (VA)

On-site
USD 86,000 - 138,000
Incident manager
Incident manager

RxCloud • El Segundo (CA)

On-site
USD 100,000 - 130,000
Technology Incident Manager
Technology Incident Manager

Burtch Works • Brooklyn (OH)

On-site
USD 90,000 - 130,000
Urgent Client Requirement: Incident & Request Manager 19997-1
Urgent Client Requirement: Incident & Request Manager 19997-1

Jobs via Dice • Atlanta (GA)

On-site
USD 100,000 - 130,000
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • Herndon (VA)

On-site
USD 86,000 - 138,000
EOC Incident Manager - Watch Officer
EOC Incident Manager - Watch Officer

Dunhill Professional Search & Government Solutions • Ashburn (VA)

On-site
USD 110,000 - 150,000
Incident Manager - Senior [Customer IT Support]
Incident Manager - Senior [Customer IT Support]

platacard • United States

Hybrid
USD 150,000 - 190,000
Relocation support to Mexico
Healthcare coverage
Education budget for language lessons,
+2
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • United States

On-site
USD 86,000 - 138,000
Medical and dental insurance
401(k) plan
Paid time off (PTO)