Lead Incident Management

Tata Consultancy Services

Chicago (IL)

On-site

USD 100,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Incentive
Medical Coverage
Family Leaves
401K Plan
Training Reimbursement
Vacation & Holidays

Job summary

Tata Consultancy Services in the United States is seeking an experienced Major Incident Manager to own and coordinate P1 and P2 incidents across Azure and on‑premises services. The role requires leadership across technical resolver teams, governance of incident playbooks, and executive communications during business‑critical outages.

You will drive rapid triage, escalation decisions, and post‑incident reviews while partnering with Change Management and Problem Management to reduce repeat

Qualifications

  • 6+ years of IT Service Management experience with a dedicated Major Incident focus.
  • Experience as Incident Commander in a large enterprise at Fortune 500/FTSE 100 complexity.
  • ITIL 4 Managing Professional or High Velocity IT certification (Foundation minimum).
  • Hands-on Azure incident management: Monitor, Service Health, Log Analytics, Application Insights, MS support escalations.

Responsibilities

  • Lead Major Incident Command & Coordination across on-premises and Azure services.
  • Serve as the single owner for all P1/P2 incidents from declaration to closure.
  • Chair live incident bridges and rotate through resolver groups and partners.
  • Drive rapid triage using Azure dashboards to define scope and blast radius within 15 minutes.
  • Maintain stakeholder communications, including executive briefings during outages.
  • Facilitate PIRs within SLAs and drive continuous improvement.

Skills

ITSM
Major Incident Command
Azure Monitor
Azure Service Health
Log Analytics
Application Insights
ServiceNow ITSM
KPI Dashboards
Post-Incident Reviews
Stakeholder Communications
Incident KPI Metrics

Tools

ServiceNow ITSM
Microsoft Teams
Azure Monitor
Log Analytics

Job description

Job Description
  • 6+ years of IT Service Management experience with a minimum of 3 years in a dedicated Major Incident
  • Management or Incident Commander role in a large enterprise (Fortune 500 / FTSE 100 equivalent complexity).
  • ITIL 4 Managing Professional or ITIL 4 Specialist: High Velocity IT certification
  • Demonstrable experience managing Azure platform incidents: working knowledge of Azure Monitor
  • Azure Service Health, Log Analytics, Application Insights, and Microsoft support escalation paths.
  • Proven ability to command high-pressure P1 incidents involving 20+ stakeholders across technical and executive levels simultaneously
  • Expert-level proficiency in ServiceNow ITSM, including Incident, Problem, Change modules and
  • Strong data analysis skills: ability to analyze incident trends, build KPI dashboards, and present actionable insights to senior leadership.
  • Major Incident Command & Coordination
  • Serve as the single accountable owner for all P1 and P2 major incidents across on premises and Azure-hosted services, from initial declaration through resolution and post-incident closure.
  • Convene and chair live incident bridge calls and virtual war rooms using Microsoft Teams, coordinating across 10+ internal technical resolver groups, managed service partners
  • Drive swift triage by leveraging Azure Service Health, Resource Health, and Azure Monitor dashboards to rapidly establish scope, affected services, and blast radius within the first 15 minutes of an incident.
  • Make and enforce escalation decisions, including engaging Microsoft CSS P1 Severity A support cases and activating DR runbooks where service restoration via normal means is not achievable within RTO.
  • Maintain clear, timely, and audience-appropriate stakeholder communications throughout the incident lifecycle, including CEO/CISO executive briefings for business-critical outages.
  • Facilitate structured blameless Post-Incident Reviews (PIRs) within agreed SLAs (P1: 48 hours, P2: 5 business days); produce high-quality PIR reports consumed by CTO and Board Technology Committee.
  • Own the incident action item registry; chair weekly SIP (Service Improvement Plan) reviews to ensure commitments are delivered on time and to quality.
  • Identify systemic incident patterns through trend analysis using ServiceNow and Log Analytics.
  • Collaborate with Problem Management to drive root cause elimination for repeat incidents.
  • Define, track, and report on enterprise incident management KPIs: MTTD, MTTR, incident recurrence rate, SLA compliance, and customer impact hours — presented to IT leadership in monthly operational reviews.
  • Process Ownership & ITSM Governance
  • Own, maintain, and continuously improve the enterprise Major Incident Management process, policy, playbooks, and runbooks aligned to ITIL 4 and the organization’s IT Risk and Control Framework.
  • Define and govern the incident severity classification matrix and escalation decision tree.
  • Ensure consistent adoption across all IT towers and managed service partners.
  • Maintain and test the enterprise crisis communication framework, including stakeholder notification trees, bridge protocols, and executive communication templates.
  • Collaborate with Change Management to ensure CAB processes adequately assess change-induced incident risk; maintain correlation tracking between changes and incidents.
Must Have Technical/Functional Skills
  • 6+ years of IT Service Management experience with a minimum of 3 years in a dedicated Major Incident
  • Management or Incident Commander role in a large enterprise (Fortune 500 / FTSE 100 equivalent complexity).
  • ITIL 4 Managing Professional or ITIL 4 Specialist: High Velocity IT certification (ITIL 4 Foundation minimum required).
  • Demonstrable experience managing Azure platform incidents: working knowledge of Azure Monitor, Azure Service Health, Log Analytics, Application Insights, and Microsoft support escalation paths.
  • Proven ability to command high-pressure P1 incidents involving 20+ stakeholders across technical and executive levels simultaneously
  • Expert-level proficiency in ServiceNow ITSM, including Incident, Problem, Change modules and dashboard/report building.
  • Strong data analysis skills: ability to analyze incident trends, build KPI dashboards, and present actionable insights to senior leadership.
Roles & Responsibilities
  • Major Incident Command & Coordination
  • Serve as the single accountable owner for all P1 and P2 major incidents across on premises and Azure-hosted services, from initial declaration through resolution and post-incident closure.
  • Convene and chair live incident bridge calls and virtual war rooms using Microsoft Teams, coordinating across 10+ internal technical resolver groups, managed service partners, and Microsoft Azure Support (Unified Support escalations).
  • Drive swift triage by leveraging Azure Service Health, Resource Health, and Azure Monitor dashboards to rapidly establish scope, affected services, and blast radius within the first 15 minutes of an incident.
  • Make and enforce escalation decisions, including engaging Microsoft CSS P1 Severity A support cases and activating DR runbooks where service restoration via normal means is not achievable within RTO.
  • Maintain clear, timely, and audience-appropriate stakeholder communications throughout the incident lifecycle, including CEO/CISO executive briefings for business-critical outages.
Post-Incident Review & Continual Improvement
  • Facilitate structured blameless Post-Incident Reviews (PIRs) within agreed SLAs (P1: 48 hours, P2: 5 business days); produce high-quality PIR reports consumed by CTO and Board Technology Committee.
  • Own the incident action item registry; chair weekly SIP (Service Improvement Plan) reviews to ensure commitments are delivered on time and to quality.
  • Identify systemic incident patterns through trend analysis using ServiceNow and Log Analytics.
  • Collaborate with Problem Management to drive root cause elimination for repeat incidents.
  • Define, track, and report on enterprise incident management KPIs: MTTD, MTTR, incident recurrence rate, SLA compliance, and customer impact hours — presented to IT leadership in monthly operational reviews.
Azure Operations & Cloud Incident Specifics
  • Develop and maintain Azure-specific incident playbooks covering platform scenarios: AKS node/pod failures, Azure SQL failover events, ExpressRoute circuit drops, Azure Active Directory (Entra ID) authentication outages, and Azure region-wide service incidents.
  • Maintain working relationships with Microsoft TAM (Technical Account Manager) and Azure Rapid Response team: ensure escalation paths to Microsoft CSS are exercised and SLAs understood.
  • Monitor Azure Service Health and Microsoft 365 Service Health Dashboard proactively.
  • Initiate pre-emptive incident declarations for advisory/degraded-service notifications affecting business-critical services.
  • Participate in Azure Operational Reviews with Cloud Platform and SRE teams to identify observability gaps, alerting blind spots, and runbook deficiencies before they manifest as major incidents.
Capability Building & Stakeholder Engagement
  • Design and deliver MIM process training programmes for Level 1/2 Service Desk, resolver groups, and technology leadership; conduct quarterly simulation exercises (GameDay / IncidentEx).
  • Act as a subject matter expert in enterprise-wide DR and BCP exercises; validate incident response readiness across all Azure-hosted Tier-0 services.
  • Build and manage a network of Incident Coordinators across global IT towers to support follow-the-sun incident coverage.

Salary Range: $100,000-$120,000 a year

TCS Employee Benefits Summary
  • Discretionary Annual Incentive.
  • Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
  • Family Support: Maternal & Parental Leaves.
  • Insurance Options: Auto & Home Insurance, Identity Theft Protection.
  • Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
  • Time Off: Vacation, Time Off, Sick Leave & Holidays.
  • Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Incident Management
Lead Incident Management

Tata Consultancy Services • Deerfield (IL)

On-site
USD 100,000 - 120,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+3
System Administrator
System Administrator

Tata Consultancy Services • Bellevue (WA)

On-site
USD 100,000 - 125,000
Discretionary Annual Incentive
Medical Coverage: Medical & Health, 4D
Dental & Vision
+4
Service Delivery Manager
Service Delivery Manager

Tata Consultancy Services • New York (NY)

On-site
USD 100,000 - 140,000
Discretionary Incentive
Medical Coverage
Family Leaves
+4
Cloud & Infrastructure Engineering / Architect
Cloud & Infrastructure Engineering / Architect

Tata Consultancy Services • Richfield (MN)

On-site
USD 120,000 - 145,000
Discretionary Annual Incentive
Comprehensive Medical Coverage: Health
Family Support: Leaves
+3
DR Coordinator
DR Coordinator

Tata Consultancy Services • Newark (NJ)

On-site
USD 100,000 - 110,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+1
Data Centre operation support
Data Centre operation support

Tata Consultancy Services • BLOOMINGTON (MN)

On-site
USD 60,000 - 70,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental, Vis,
Parental Leaves
+1
Azure Cloud Engineer II
Azure Cloud Engineer II

iT1 • Tempe (AZ)

On-site
USD 90,000 - 120,000
Medical, dental, and vision benefits
Paid time off
401(k) Plan with employer match
+2
Data Centre operation support
Data Centre operation support

Tata Consultancy Services • Bloomington (IL)

On-site
USD 60,000 - 70,000
Discretionary incentive
Comprehensive medical coverage
Family support: parental leaves
+4
Azure Cloud Engineer II
Azure Cloud Engineer II

It1 Source LLC • Tempe (AZ)

On-site
USD 90,000 - 120,000
Medical, dental, and vision benefits
401(k) Plan with employer match
Onsite Fitness Center
Director of Platform Engineering & Operations, Hands-on, Onsite in Charlotte, NC
Director of Platform Engineering & Operations, Hands-on, Onsite in Charlotte, NC

Gina’s Tech Jobs - IT Recruiting Agency • Charlotte (NC)

On-site
USD 130,000 - 160,000
Medical insurance
Dental
Vision
+2