Technology Support Lead - Problem Management & Governance Lead

JPMorgan Chase & Co.

Columbus (OH)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorgan Chase & Co. seeks a Problem Management Lead in Employee Platforms to serve as SME for end-to-end problem management, focusing on root cause identification, reduced incidents, and stable services. The role emphasizes risk and control standards with data-driven recommendations.

You will drive remediation actions, coordinate with engineering, and ensure remediation aligns with resiliency and security expectations, while reporting progress to stakeholders.

Qualifications

  • 5+ years of experience troubleshooting, resolving, and maintaining IT services.
  • Experience using enterprise-authorized AI capabilities to support production operations workflows.
  • Ability to review and validate AI-assisted incident recommendations before action.
  • Experience managing applications or infrastructure in large-scale environments (on-premises and cloud).
  • Proficient in observability and monitoring tools and techniques.
  • Experience executing ITIL-based processes.

Responsibilities

  • Serve as SME for end-to-end problem management across core services from identification to closure.
  • Lead adoption of enterprise AI capabilities to improve incident triage speed and consistency.
  • Manage incident, problem, and change management for full-stack systems.
  • Analyze incident trends to identify systemic problems and prioritize remediation.
  • Facilitate root cause analyses and post-incident reviews, documenting corrective actions.
  • Establish ownership, milestones, and reporting for problem records and remediation.

Skills

Problem management
AI in production ops
Root cause analysis
Stakeholder management
Observability/monitoring
ITIL familiarity

Education

ITIL certification

Job description

Join our dynamic Production Management team to strengthen technology resilience, improve service quality, and drive effective governance across the core services that support our business.

As a Problem Management Lead in Employee Platforms, you will serve as a subject matter expert for end-to-end problem management, with a primary focus on identifying root causes, reducing recurring incidents, improving service stability, and ensuring that remediation and closure decisions meet established risk, control, resiliency, and documentation standards.

Success in this role requires strong analytical thinking, sound judgment, attention to detail, and the ability to influence technology teams and business stakeholders through clear communication, effective challenge, and data-driven recommendations.

Job responsibilities

  • Serve as the subject matter expert for end-to-end problem management across application and infrastructure services, from problem identification and classification through root cause analysis, remediation, validation, and closure.

  • Leads team adoption of enterprise-authorized AI capabilities within the work environment to improve incident triage speed and consistency (e.g., synthesizing operational signals into prioritized actions), with human-in-the-loop validation and appropriate handling of sensitive data.

  • Execute policies and procedures that ensure operational stability and availability and lead incident, problem, and change management in support of full stack technology systems, applications, or infrastructure.

  • Analyze incident trends, recurring issues, service-impacting events, and operational data to identify systemic problems and prioritize remediation based on business impact, risk, and customer experience.

  • Facilitate structured root cause analyses and post-incident reviews, ensuring contributing factors, control gaps, and lessons learned are documented and translated into sustainable corrective and preventive actions.

  • Establish clear ownership, milestones, and success measures for problem records and remediation actions; monitor progress, escalate delays, and provide transparent reporting to technology and business stakeholders.

  • Partner with engineering, operations, service owners, and control functions to deliver sustainable remediation, improve service resilience, and reduce the recurrence and business impact of known issues.

  • Maintain problem records and known-error information, ensuring workarounds, technical findings, risk acceptances, and remediation plans remain current, accessible, and aligned with established standards.

  • Identify operational risks and control deficiencies through problem management activities and coordinate appropriate mitigation, governance, audit, and compliance actions.

  • Define and report problem management metrics, including recurrence, aging, remediation effectiveness, root cause categories, and reductions in incident volume or business impact.

  • Applies reuse-first, AI-assisted practices across incident/problem/change routines to identify recurring interruption patterns and validate remediation actions aligned to resiliency and security expectations.

Required qualifications, capabilities, and skills

  • 5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services.

  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows with strong validation habits and awareness of data sensitivity.

  • Ability to review and validate AI-assisted incident recommendations before action, escalating when uncertain and ensuring outcomes align to operational, security, and auditability expectations.

  • Experience managing applications or infrastructure in a large-scale technology environment both on premises and public cloud.

  • Proficient in observability and monitoring tools and techniques.

  • Experience executing on processes in scope of the Information Technology Infrastructure Library (ITIL) framework.

  • Demonstrated experience managing end-to-end problem records, root cause analyses, corrective actions, known errors, and recurrence-prevention activities.

  • Strong analytical skills with experience using incident trends and operational data to identify systemic problems and prioritize remediation based on impact and risk.

  • Ability to manage operational risk, control requirements, resiliency expectations, and auditability needs within a complex technology environment.

  • Proven stakeholder management and influencing skills, including the ability to challenge constructively and communicate effectively with senior business and technology stakeholders.

Preferred qualifications, capabilities, and skills

  • Experience supporting Problem Management, Production Management, Site Reliability Engineering, or Operational Excellence activities within a large enterprise environment.

  • Knowledge of structured root cause analysis methods, post-incident review practices, known-error management, and corrective and preventive action management.

  • Experience defining and reporting problem management metrics, including recurrence, aging, remediation effectiveness, root cause categories, and service impact reduction.

  • Familiarity with enterprise-authorized AI, automation, analytics, or AIOps capabilities used to support operational analysis and workflow efficiency.

  • Relevant industry certification or formal training in ITIL, Service Management, Site Reliability Engineering, cloud technology, or operational risk management.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technology Support Lead - AI and Production Management Tools
Technology Support Lead - AI and Production Management Tools

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 180,000 - 240,000
Technology Support Lead - Major Incident Management
Technology Support Lead - Major Incident Management

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 110,000 - 170,000
Technology Support III - Incident Manager
Technology Support III - Incident Manager

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 70,000 - 100,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
Incident and Problem Management Analyst
Incident and Problem Management Analyst

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 100,000 - 130,000
Operations Specialist
Operations Specialist

Aquent • Southlake (TX)

On-site
USD 75,000 - 110,000
ITSM Problem Manager
ITSM Problem Manager

NextGen Healthcare • United States

On-site
USD 90,000 - 120,000
Program Manager
Program Manager

Compunnel, Inc. • Worcester (MA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Application Engineer View role →
Application Engineer View role →

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Production Support Engineer
Production Support Engineer

ASM Tech Solutions • United States

Hybrid
USD 90,000 - 150,000