Staff Infrastructure Engineer

The Hartford

Charlotte (NC)

Hybrid

USD 117,000 - 175,000

Full time

27 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

The Hartford is seeking a Staff Infrastructure Engineer to leverage AI-driven methods for rapid detection, assessment, and resolution of critical service issues and major incidents. You will work within the Technology Command Center to enhance enterprise resilience and implement proactive AI-driven operations.

Hybrid work schedule is expected, with 3 days in the office (Tue-Thu) at Hartford, CT or Charlotte, NC, and cross-functional collaboration with infrastructure, reliability engineering,

Qualifications

  • 8+ years' experience in monitoring and observability tools such as Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes, or similar.
  • Experience in AI-driven data analysis and problem solving.
  • AI tool proficiency including prompt engineering and Interrogation skills
  • Familiarity with ITSM and ticketing platforms, including incident creation, categorization, escalation, and documentation.
  • Understanding of event correlation, alert prioritization, service impact analysis, and basic automation/workflow enablement.
  • Strong analytical thinking and operational judgment.
  • Ability to work under pressure during high-impact events.

Responsibilities

  • Manage technical and executive communication for major incidents.
  • Perform alert triage, correlation, and initial impact assessment to separate actionable events from non-actionable noise.
  • Support major incident detection and escalation by validating symptoms, confirming affected services, and engaging resolver teams.
  • Use standard operating procedures, runbooks, and decision frameworks to investigate, prioritize, and escalation events.
  • Maintain situational awareness during active events; document event patterns and operational observations.
  • Promote events to incidents when thresholds are met; partner with internal and external teams during restoration activities.
  • Identify opportunities to improve monitoring effectiveness, event quality, automation, and operational readiness.

Skills

Splunk
Dynatrace
ITSI
Moogsoft
ThousandEyes
AI analytics
Incident management
Prompt engineering
Automation

Tools

Splunk
Dynatrace
ITSM tools
Moogsoft
ThousandEyes

Job description

Staff Infrastructure Engineer - IEB07CE

We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future.

The Staff Infrastructure Engineer is responsible for leveraging agentic AI to quickly identify, assess, and resolve critical service issues and major incidents, ensuring swift detection and response. The successful candidate will demonstrate engineering-level technical depth in AI, cloud computing concepts, and infrastructure platforms, partnering with cross-functional teams for service restoration and continuously improving the event management process. This position is critical in enhancing enterprise resilience and implementing proactive AI-driven operations within the Technology Command Center.

This role will have a Hybrid work schedule, with the expectation of working in an office (Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).

Key Responsibilities
  • Manage technical and executive communication for major incidents.
  • Perform alert triage, correlation, and initial impact assessment to separate actionable events from non-actionable noise.
  • Support major incident detection and escalation by validating symptoms, confirming affected services, and engaging resolver teams.
  • Use standard operating procedures, runbooks, and decision frameworks to investigate, prioritize, and escalation events.
  • Maintain situational awareness during active events; document event patterns and operational observations.
  • Promote events to incidents when thresholds are met; partner with internal and external teams during restoration activities.
  • Identify opportunities to improve monitoring effectiveness, event quality, automation, and operational readiness.
Partnership & Work Expectations
  • Collaborate with infrastructure, application, reliability engineering, service desk, and vendor teams to ensure effective event response and service restoration.
  • Support continuous improvement by identifying recurring alert issues, documenting operational insights, and recommending enhancements to monitoring, runbooks, and automation.
Schedule & Role Expectations
  • May require shift coverage, off-hours support, and escalation activities aligned to a 24x7 operational model and follow-the-sun support structure.
Required Qualifications
  • 8+ years' experience in monitoring and observability tools such as Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes, or similar.
  • Experience in AI-driven data analysis and problem solving.
  • AI tool proficiency including prompt engineering and Interrogation skills
  • Familiarity with ITSM and ticketing platforms, including incident creation, categorization, escalation, and documentation.
  • Understanding of event correlation, alert prioritization, service impact analysis, and basic automation/workflow enablement.
  • Strong analytical thinking and operational judgment.
  • Ability to work under pressure during high-impact events.
Candidate must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.
Compensation

The listed annualized base pay range is primarily based on analysis of similar positions in the external market. Actual base pay could vary and may be above or below the listed range based on factors including but not limited to performance, proficiency and demonstration of competencies required for the role. The base pay is just one component of The Hartford’s total compensation package for employees. Other rewards may include short-term or annual bonuses, long-term incentives, and on-the-spot recognition. The annualized base pay range for this role is:

$116,800 - $175,200

Equal Opportunity Employer/Sex/Race/Color/Veterans/Disability/Sexual Orientation/Gender Identity or Expression/Religion/Age

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Infrastructure Engineer
Staff Infrastructure Engineer

thehartford • Hartford (CT)

Hybrid
USD 117,000 - 175,000
Staff Infrastructure Engineer
Staff Infrastructure Engineer

The Hartford • Hartford (CT), Northern (KY)

On-site
USD 117,000 - 175,000
Staff Software Engineer
Staff Software Engineer

The Hartford • Columbus (OH)

On-site
USD 116,000 - 174,000
Staff Software Engineer
Staff Software Engineer

The Hartford • Chicago (IL)

On-site
USD 116,000 - 174,000
Staff Software Engineer
Staff Software Engineer

The Hartford • Charlotte (NC)

Hybrid
USD 116,000 - 174,000
Staff Software Engineer
Staff Software Engineer

The Hartford • Hartford (CT)

Hybrid
USD 116,000 - 174,000
Staff Infrastructure Engineer: AI-Driven Incident Response
Staff Infrastructure Engineer: AI-Driven Incident Response

The Hartford • Charlotte (NC)

Hybrid
USD 117,000 - 175,000
Senior Staff Platform Engineer
Senior Staff Platform Engineer

thehartford • Hartford (CT)

On-site
USD 137,000 - 206,000
Staff Software Engineer
Staff Software Engineer

The Hartford • Charlotte (NC)

On-site
USD 114,720 - 172,080
Staff Infrastructure Engineer: AI-Driven Incidents (Hybrid)
Staff Infrastructure Engineer: AI-Driven Incidents (Hybrid)

thehartford • Hartford (CT)

Hybrid
USD 117,000 - 175,000