Incident Response Lead - Production Reliability & Automation

UST

Richmond (VA)

Hybrid

USD 83,000 - 125,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with employer matching
Paid time off
Paid holidays
Life insurance

Job summary

UST is seeking an Incident Manager to handle escalated items from members and the Business concierge, coordinating with downstream teams (Application, API, EDP, L2/L3) to drive resolution. You will analyze incidents end-to-end for the Portals, monitor logs with Splunk and Datadog, and drive weekly trend reports to help prevent future outages.

Responsibilities include triaging critical production issues, coordinating with multiple tech teams, and ensuring recovery within SLA while addressing

Qualifications

  • Experience analyzing incidents and coordinating with multiple teams.
  • Strong knowledge of monitoring tools and incident management processes.
  • Excellent communication across leadership levels and functions.

Responsibilities

  • End-to-end incident analysis for Portals and escalations.
  • Coordinate MTM JIRA requests with Anthem audit teams.
  • Analyze logs in Splunk and Datadog to identify trends.
  • Produce weekly incident trend reports and aging metrics.
  • Bridge gaps between members and development teams on ad-hoc requests.
  • Support bulk data uploads and collaborate with Product teams.
  • Identify automation opportunities and address privacy/compliance concerns.
  • Triage critical production issues and coordinate recovery within SLA.

Skills

Splunk
Java API analysis
Apigee
Microservices
HTTP
OAuth
SSL
Java
SNOW (ServiceNow)
Datadog
AWS
Incident Analysis and management
Web services monitoring
API monitoring
Excellent oral and written comms
Clear articulation (verbal/written)
Leadership communication about tech
Triage and diagnose performance issues

Tools

Splunk
Apigee
Datadog
SNOW (ServiceNow)
AWS

Job description

UST is seeking an Incident Manager to handle escalated items from members and the Business concierge, coordinating with downstream teams (Application, API, EDP, L2/L3) to drive resolution. You will analyze incidents end-to-end for the Portals, monitor logs with Splunk and Datadog, and drive weekly trend reports to help prevent future outages.

Responsibilities include triaging critical production issues, coordinating with multiple tech teams, and ensuring recovery within SLA while addressing

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Incident Command Lead – Production Reliability
Incident Command Lead – Production Reliability

State Street • Boston (MA)

On-site
USD 70,000 - 118,750
Production Incident & Reliability Lead
Production Incident & Reliability Lead

State Street • Boston (MA)

On-site
USD 70,000 - 119,000
401K with company match
Healthcare coverage (medical, dental,
Paid time off and disability benefits
+3
Production Services Specialist with Incident Management
Production Services Specialist with Incident Management

KPG99 INC • Chandler (AZ)

On-site
USD 90,000 - 150,000
Senior Production Support Lead: Incident & Reliability
Senior Production Support Lead: Incident & Reliability

STATE STREET CORPORATION • Boston (MA)

On-site
USD 70,000 - 119,000
401K with company match
Health insurance
Paid time off
+1
Incident Response Lead: Tech Outages & Readiness
Incident Response Lead: Tech Outages & Readiness

Horizontal Talent • Brooklyn (OH)

On-site
USD 69,000 - 77,000
Medical insurance
Dental insurance
Vision insurance
+1
Senior Incident & Escalations Lead
Senior Incident & Escalations Lead

Outreach • Atlanta (GA)

Hybrid
USD 95,000 - 145,000
Production Operations & Incident Analyst
Production Operations & Incident Analyst

UZURV – The Adaptive TNC • Richmond (VA)

On-site
USD 65,000 - 90,000
401K matching
Healthcare benefits package
Generous PTO and paid holidays
+1
Incident Response Lead — Customer Advocate & Orchestrator
Incident Response Lead — Customer Advocate & Orchestrator

Staffingine LLC • Redmond (WA)

On-site
USD 110,000 - 150,000
Production Operations Lead - Incident & Reliability
Production Operations Lead - Incident & Reliability

Compunnel, Inc. • Worcester (MA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Incident Leader – 24/7 Enterprise Response
Senior Incident Leader – 24/7 Enterprise Response

System One • Knoxville (TN)

On-site
USD 110,000 - 160,000