Problem Manager – ITIL, ServiceNow, Root Cause Analysis

Jobtailor

Utah

On-site

USD 90,000 - 120,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced Problem Management professional to lead end-to-end lifecycles from investigation through closure. You will collaborate with IT and business teams to drive remediation action plans, ensure root-cause solutions, and prevent recurrence across multi-domain environments.

Responsibilities include facilitating 5 Whys RCA sessions, analyzing logs and dashboards, and leveraging AI-assisted tools for timelines and RCA documentation.

Qualifications

  • 7+ years of experience in IT Operations, Problem Management, Incident Management, Major Incident Management, or related disciplines.
  • Strong understanding of ITIL Problem Management, Incident Management, and operational governance.
  • Experience conducting Root Cause Analysis, 5 Whys investigations, and post-incident reviews.
  • Hands-on experience with ServiceNow or comparable ITSM platforms.
  • Knowledge of Windows and Linux platforms.
  • Knowledge of VMware and Hyper-V virtualization.

Responsibilities

  • Facilitate timely 5 Whys Root Cause Analysis sessions for P1 incidents and recurring P2 incidents
  • Review system logs, monitoring tools, dashboards, and change records to establish timelines and contributing factors
  • Use AI-assisted tools for incident investigations, timeline generation, RCA documentation, and knowledge capture
  • Document root causes and corrective actions in ServiceNow Problem Management records
  • Lead the end-to-end Problem Management lifecycle from investigation through closure
  • Partner with technical teams to create Remediation Action Plans
  • Track remediation, accelerate risks, and hold stakeholders accountable
  • Validate solutions address root causes and prevent recurrence
  • Measure and report service reliability, operational risk, and incident recurrence improvements
  • Identify recurring patterns across infrastructure, cloud, application, and network incidents
  • Drive proactive problem identification and long-term service improvements
  • Recommend monitoring enhancements, automation opportunities, and architectural improvements
  • Facilitate weekly Problem Review meetings
  • Maintain executive dashboards and Problem Management reporting
  • Deliver executive-ready summaries and status updates
  • Produce RCA reports for customers and senior leadership
  • Improve governance processes, operational standards, runbooks, and reporting practices
  • Champion automation and AI-enabled workflows

Skills

Analytical skills
Documentation skills
Facilitation skills
Stakeholder management
Executive communications

Education

ITIL Foundation

Tools

ServiceNow
Monitoring Tools
Dashboards
AI-Assisted Tools

Job description


  • Facilitate timely 5 Whys Root Cause Analysis sessions for P1 incidents and recurring P2 incidents

  • Review system logs, monitoring tools, dashboards, and change records to establish timelines and contributing factors

  • Use AI-assisted tools for incident investigations, timeline generation, RCA documentation, and knowledge capture

  • Document root causes and corrective actions in ServiceNow Problem Management records

  • Lead the end-to-end Problem Management lifecycle from investigation through closure

  • Partner with technical teams to create Remediation Action Plans

  • Track remediation, accelerate risks, and hold stakeholders accountable

  • Validate solutions address root causes and prevent recurrence

  • Measure and report service reliability, operational risk, and incident recurrence improvements

  • Identify recurring patterns across infrastructure, cloud, application, and network incidents

  • Drive proactive problem identification and long-term service improvements

  • Recommend monitoring enhancements, automation opportunities, and architectural improvements

  • Facilitate weekly Problem Review meetings

  • Maintain executive dashboards and Problem Management reporting

  • Deliver executive-ready summaries and status updates

  • Produce RCA reports for customers and senior leadership

  • Improve governance processes, operational standards, runbooks, and reporting practices

  • Champion automation and AI-enabled workflows


Requirements


  • 7+ years of experience in IT Operations, Problem Management, Incident Management, Major Incident Management, or related disciplines

  • Strong understanding of ITIL Problem Management, Incident Management, and operational governance

  • Experience conducting Root Cause Analysis, 5 Whys investigations, and post-incident reviews

  • Hands-on experience with ServiceNow or comparable ITSM platforms

  • Strong analytical, documentation, facilitation, and stakeholder management skills

  • Ability to influence cross-functional teams and drive accountability without direct authority

  • Experience creating executive-level communications, dashboards, and operational reporting

  • Broad technical understanding across enterprise technology environments

  • Knowledge of Windows and Linux platforms

  • Knowledge of VMware and Hyper-V virtualization

  • Understanding of performance troubleshooting for CPU, memory, storage, and I/O

  • Knowledge of SAN and NAS environments and storage performance, redundancy, and resiliency concepts

  • Knowledge of routing, switching, VLANs, DNS, firewalls, load balancing, packet flow, latency, and packet loss

  • Knowledge of AWS and Microsoft Azure, including cloud networking, compute, identity, and storage

  • Understanding of APIs, microservices, distributed systems, CI/CD, deployment pipelines, application defects, configuration drift, and dependencies

  • ITIL Foundation or higher certification preferred

  • Experience supporting large-scale enterprise environments preferred

  • Experience with operational analytics, trend analysis, and KPI reporting preferred

  • Exposure to automation, AI-enabled workflows, or AIOps preferred

  • Experience with executive stakeholders and client-facing incident communications preferred


Core Competencies

Demonstrates expertise in ITIL Problem Management and Incident Management, with a strong focus on conducting Root Cause Analysis and driving service reliability improvements. Proficient in utilizing ServiceNow for documentation and reporting, while effectively communicating with executive stakeholders.


Highest-signal resume keywords


  • ITIL Problem Management

  • Root Cause Analysis

  • ServiceNow

  • Operational Reporting

  • Cloud Networking


ATS Optimization Keywords

Hard Skills


  • 5 Whys Investigation

  • Operational Analytics

  • Performance Troubleshooting

  • API Understanding

  • CI/CD

  • Virtualization (VMware, Hyper-V)

  • Storage Performance (SAN, NAS)

  • Networking (Routing, Switching, VLANs)

  • Cloud Platforms (AWS, Microsoft Azure)

  • Documentation


Soft Skills


  • Analytical Skills

  • Facilitation Skills

  • Stakeholder Management

  • Influencing Skills

  • Communication Skills


Certifications & Qualifications


  • ITIL Foundation


Industry Keywords


  • IT Operations

  • Incident Management

  • Major Incident Management

  • Operational Governance

  • Service Reliability


Tools & Technologies


  • ServiceNow

  • Monitoring Tools

  • Dashboards

  • AI-Assisted Tools

  • Executive Dashboards

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Service Desk Supervisor
Lead Service Desk Supervisor

Jobtailor • Arizona

On-site
USD 90,000 - 130,000
Service Reliability Specialist
Service Reliability Specialist

Jobtailor • Town of Florida (NY)

On-site
USD 90,000 - 130,000
Director of Service Operations
Director of Service Operations

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Enterprise SRE – Incident Coordinator
Enterprise SRE – Incident Coordinator

Jobtailor • Missouri

On-site
USD 90,000 - 120,000
Technical Platform Operations Support, Manager
Technical Platform Operations Support, Manager

Jobtailor • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Senior Systems Engineer, AI-Enabled Service Delivery
Senior Systems Engineer, AI-Enabled Service Delivery

Jobtailor • La Verne (CA)

Hybrid
USD 80,000 - 110,000
Senior Manager, IT Service Management – Product Lead
Senior Manager, IT Service Management – Product Lead

Jobtailor • Connecticut

On-site
USD 120,000 - 150,000
Senior Windows Systems Engineer, Admin
Senior Windows Systems Engineer, Admin

Jobtailor • Phillipsburg (NJ)

On-site
USD 120,000 - 180,000
IT ServiceNow Engineer
IT ServiceNow Engineer

Jobtailor • Colorado

On-site
USD 110,000 - 170,000
Customer Delivery Manager
Customer Delivery Manager

Jobtailor • California (MO)

On-site
USD 120,000 - 170,000