Technical Enterprise Incident Manager

Peraton

Linthicum (MD)

Remote

USD 120,000 - 160,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Peraton is seeking a Technical Enterprise Incident Manager to lead enterprise incident response, service restoration, and reliability programs. You’ll coordinate across infrastructure, cloud, applications, and vendors from a central point of contact, ensuring rapid resolution and clear communications to leadership.

The role emphasizes hands-on cloud and infrastructure expertise, ITIL/SRE practices, and ongoing improvements to monitoring, runbooks, and post-incident reviews in a remote,

Qualifications

  • Must be a U.S. citizen with ability to obtain Public Trust clearance.
  • Bachelor’s Degree + 5 years of experience or HS diploma + 9 years.
  • 5+ years in Cloud Incident Management, Operations Engineering, NOC, SRE or production support.
  • Experience leading enterprise Major Incident response in 24x7 operations.
  • Strong ITIL Incident and Problem Management understanding.
  • 3+ years with AWS cloud services; familiarity with monitoring/observability tools.
  • Experience with ServiceNow or similar ITSM platforms; strong communication.

Responsibilities

  • Lead incident bridges across infra, apps, network, cloud, security, and vendors.
  • Drive rapid service restoration with clear timelines and executive updates.
  • Prioritize incidents by business impact and risk; manage escalations.
  • Monitor SLA compliance and lead PIRs with RCA follow-through.
  • Drive automation of detection, triage, and response workflows.
  • Maintain structured communication during major incidents and runbooks.
  • Validate runbooks and service maps; use observability tools for analysis.
  • Collaborate with teams to implement root cause remediation and improvements.

Skills

Incident management
ITIL processes
SRE
Cloud platforms (AWS)
Datadog
Cloudcraft
ServiceNow (ITSM)
Executive communication

Education

Bachelor's Degree
High School diploma or equivalent

Tools

AWS
Datadog
Cloudcraft
ServiceNow

Job description

Basic Qualifications:
  • Must be a U.S. citizen with the ability to obtain and maintain the required Public Trust level clearance
  • Bachelor’s Degree and 5 years of experience, or a High School diploma or equivalent and 9 years of experience
  • 5+ years of experience in Cloud Incident Management, Operations Engineering, NOC, SRE, Application or Production Support environments.
  • Experience leading enterprise Major Incident response efforts in a 24x7 operational environment.
  • Strong understanding of ITIL Incident and Problem Management processes.
  • 3+ years of experience working with AWS cloud services
  • Experience with monitoring and observability platforms such as Datadog, Cloudcraft, or similar
  • Experience using ServiceNow or similar ITSM platforms.
  • Strong analytical, troubleshooting, and organizational skills.
  • Excellent written and verbal communication skills with ability to facility meetings as well as brief technical teams and executive leadership.
Preferred Qualifications:
  • Experience in a Site Reliability Engineering (SRE) or Cloud Platform DevOps environment.
  • Experience supporting federal, healthcare, financial, or other highly regulated environments.
  • Hands-on experience with infrastructure technologies including:
    • Windows/Linux Servers
    • Networking concepts
    • Cloud platforms (AWS, Azure, or GCP)
    • Load balancers, proxies, DNS, and firewalls

Peraton is seeking a highly motivated and technically skilled Technical Enterprise Incident Manager with strong Cloud Platform and application experience to lead enterprise incident response, service restoration efforts, and operational reliability initiatives. This individual will serve as the central point of coordination during major incidents, ensuring rapid resolution, clear communication, and continuous service improvement across enterprise infrastructure and applications.

The ideal candidate possesses a strong operational background, excellent communication skills, and hands‑on technical expertise in infrastructure, cloud technologies, monitoring, automation, and IT service management processes. This role requires the ability to drive incident response while also identifying systemic reliability improvements.

Location: RemoteShift Schedule: 8am – 5pm Eastern Standard Time (EST). This position will also participate in 24x7 on‑call rotations for incident management.

What You Will Do
Enterprise Incident Management
  • Lead and coordinate Incident bridge calls involving infrastructure, application, network, cloud, security, and vendor teams.
  • Drive rapid service restoration while maintaining accurate timelines, communications, and executive updates.
  • Ensure incidents are prioritized appropriately based on business impact and operational risk.
  • Manage escalation procedures and engage leadership when required.
  • Monitor SLA compliance and ensure incident response metrics are consistently achieved.
  • Facilitate Post‑Incident Reviews (PIRs) and ensure high‑quality Root Cause Analysis (RCA) documentation and follow‑through on corrective actions.
  • Drive automation of incident detection, triage, and response workflows to reduce operational toil and improve resolution times.
  • Maintain structured communication frameworks during major incidents, including timely stakeholder notifications, ongoing updates, and final incident summaries.
  • Validate runbooks, service dependency maps, and technical documentation to ensure accuracy and usability during incidents.
  • Utilize additional observability tools such as log aggregation and APM to enhance troubleshooting and incident analysis.
Cloud Platform DevSecOps Engineering
  • Work with application teams to facilitate issues and implement root cause remediations.
  • Develop and enhance monitoring, alerting, and dashboarding capabilities.
  • Analyze trends, KPIs, and operational metrics to proactively identify reliability risks.
  • Support implementation of resiliency strategies including redundancy, failover, capacity planning, and performance optimization.
  • Utilize Datadog for monitoring, alert correlation, dashboards, incident investigation, and performance analysis.
  • Participate in after‑hours on‑call incident management rotation as required.
Operational Excellence
  • Develop and maintain incident management procedures, runbooks, and knowledge articles.
  • Ensure accurate ticket documentation within ServiceNow.
  • Drive continual service improvement initiatives aligned with ITIL and SRE best practices.
  • Collaborate with cross functional teams to improve communication, escalation paths, and operational workflows.
  • Support audit, compliance, and operational reporting requirements.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • Reston (VA)

On-site
USD 86,000 - 138,000
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • Herndon (VA)

On-site
USD 86,000 - 138,000
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • United States

On-site
USD 86,000 - 138,000
Medical and dental insurance
401(k) plan
Paid time off (PTO)
EOC Incident Manager - Watch Officer
EOC Incident Manager - Watch Officer

Dunhill Professional Search & Government Solutions • Ashburn (VA)

On-site
USD 110,000 - 150,000
Enterprise Incident Response Lead: Cloud & DevSecOps
Enterprise Incident Response Lead: Cloud & DevSecOps

Peraton • Herndon (VA)

On-site
USD 86,000 - 138,000
Incident Manager
Incident Manager

Insight Global • Long Beach (CA)

On-site
USD 80,000 - 100,000
Senior Incident Manager
Senior Incident Manager

System One • Knoxville (TN)

On-site
USD 110,000 - 160,000
Remote Enterprise Incident Manager – Cloud & SRE Lead
Remote Enterprise Incident Manager – Cloud & SRE Lead

Peraton • Linthicum (MD)

Remote
USD 120,000 - 160,000
Incident Management Analyst
Incident Management Analyst

Inserso • United States

On-site
USD 70,000 - 90,000
Incident Management Analyst
Incident Management Analyst

Inserso Corporation • United States

On-site
USD 70,000 - 100,000