Reliability Engineer

thehartford

Hartford (CT)

Hybrid

USD 110,000 - 170,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Hartford is seeking a Reliability Engineer to strengthen infrastructure resilience in cloud and SAAS environments. You will lead initiatives across observability, automation, and DevSecOps, collaborating with IT and engineering teams to ensure systems are stable, scalable, and secure.

The role emphasizes AI-driven tooling, proactive problem prevention, and continuous improvement of delivery and incident response processes within a hybrid work setting.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
  • 3+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.
  • Hands-on experience with observability tools: Splunk, Dynatrace, CloudWatch.
  • Deep knowledge of Infrastructure as Code (IaC) with Terraform, CloudFormation.
  • Proven ability to optimize CI/CD pipelines, automate deployments, and enforce DevSecOps best practices.
  • Expertise in cloud platforms (AWS) and Kubernetes-based microservices environments.
  • Strong proficiency in Python, Java for infrastructure automation and tooling development.
  • Experience in AI/ML frameworks for observability, predictive failure detection, and AI-driven troubleshooting desirable.
  • Experience with Oracle and SQL Server relational database technologies. Knowledge of open-source database technologies is beneficial.
  • Demonstrated experience working within Agile frameworks and methodologies.
  • Excellent analytical, problem solving and interpersonal skills.

Responsibilities

  • Assist in instrumenting code/application stacks to generate metrics on health, availability, performance and resiliency.
  • Support architecture and software engineering teams to influence technical strategy and cross-functional impact.
  • Act as a technical leader for applications supported, covering multiple technologies and domains.
  • Develop tooling, alerts, and response mechanisms to identify and address reliability risks with automation.
  • Improve delivery flow by engineering solutions that increase speed while maintaining standards.
  • Promote preventive controls and automation for cost efficiency and self-healing capabilities.
  • Drive incident triage, service restoration, and end-to-end ownership across IT operations.
  • Collaborate with infra teams to enhance monitoring/alerting and automate recovery processes.
  • Research and implement AI-based anomaly detection to predict failures and automate preventive measures.

Skills

SRE
Observability
Terraform
CloudWatch
Kubernetes
Python
Java
DevSecOps
AWS
AI/ML

Education

Bachelor's or Master's in CS/Engineering

Tools

Splunk
Dynatrace
CloudWatch
Terraform

Job description

Reliability Engineer - IE08GE

We're determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals - and to help others accomplish theirs, too. Join our team as we help shape the future.

The Hartford's Corporate / HIMCO IT team is seeking a highly motivated, detail-oriented, and results-driven Reliability Engineer to join our team. This position will play a crucial role to lead infrastructure resilience in ensuring the stability and performance of our systems in cloud and SAAS environments.

Successful candidates will be expected to demonstrate strong technical skills, excellent partnership with stakeholders and partner teams, willingness to understand existing processes and systems, solid technical acumen, experience in delivering quality technical solutions and ensure the systems are stable, performant, and secure.

Responsibilities:
  • Assist in the use of best-in-class software engineering standards and design practices for instrumenting code/application technology stack to enable the generation of relevant metrics on overall technology health - availability, performance, quality, currency and resiliency.
  • Assist the architecture and software engineering teams to influence the technical strategy for the organization, keeping in mind its cross-functional impacts, integration across the organization, and architecture rationalization.
  • Assist on a team as a technical leader for the applications supported, requiring depth and breadth of knowledge in technologies, applications, integration, interfaces and business domain.
DevSecOps Solution Responsibilities:
  • Assist in developing effective tooling, alerts, and response mechanisms to identify and address reliability risks leveraging automation to support problem prevention, detection, mitigation, and resolution.
  • Assist in enhancing the delivery flow by engineering the appropriate solutions to increase delivery speed while adhering to technology standards for sustained reliability.
  • Partner to implement preventative controls and drive increased automation and self-healing capabilities. Continue to improve cost efficiency baselines
  • Promote and implement innovative solutions.
IT Ops Responsibilities:
  • Ensure operational excellence. Collaborate to drive the triaging and service restoration of all high impact incidents in order to minimize the mean time to service restoration and impact to the business. Demonstrate end-to-end ownership.
  • Partner with infrastructure teams to design and implement intelligent incident routing, enhanced monitoring/alerting capabilities and automated service restoration processes. Take proactive measures to prevent high impactful incidents.
  • Achieve and maintain the continuity of Hartford and third-party assets that support a business function. Accountable for keeping the IT application and infrastructure metadata repositories current.
AI-Driven Automation:
  • Research and implement AI-based anomaly detection to predict infrastructure failures and automate preventive measures.
  • Develop AI-powered troubleshooting copilots and LLM-driven operational assistants to accelerate incident resolution and root cause analysis.
  • Implement AI/ML-based runbooks to automate system recovery and optimize operational efficiency.
Qualifications:
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
  • 3+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.
  • Hands-on experience with observability tools: Splunk, Dynatrace, CloudWatch.
  • Deep knowledge of Infrastructure as Code (IaC) with Terraform, CloudFormation.
  • Proven ability to optimize CI/CD pipelines, automate deployments, and enforce DevSecOps best practices.
  • Expertise in cloud platforms (AWS) and Kubernetes-based microservices environments.
  • Strong proficiency in Python, Java for infrastructure automation and tooling development.
  • Experience in AI/ML frameworks for observability, predictive failure detection, and AI-driven troubleshooting desirable.
  • Experience with Oracle and SQL Server relational database technologies. Knowledge of open-source database technologies is beneficial.
  • Demonstrated experience working within Agile frameworks and methodologies.
  • Excellent analytical, problem solving and interpersonal skills.

This role will have a Hybrid work schedule, with the expectation of working in an office (Columbus, OH, Chicago, IL, Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).

Candidates must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.

As a condition of your employment for HIMCO, you will be required to affirm to HIMCO's Code of Ethics and understand that you will be required to comply with the disclosure of accounts, holdings and pre-clearance of trades for the accounts of you and your house

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer
Reliability Engineer

The Hartford • Charlotte (NC)

Hybrid
USD 91,000 - 137,000
Reliability Engineer
Reliability Engineer

The Hartford • Columbus (OH)

Hybrid
USD 91,200 - 136,800
Reliability Engineer
Reliability Engineer

The Hartford • Chicago (IL)

Hybrid
USD 91,200 - 136,800
Hybrid work schedule
Reliability Engineer
Reliability Engineer

The Hartford • Connecticut

Hybrid
USD 91,000 - 137,000
Reliability Engineer
Reliability Engineer

The Hartford • United States

On-site
USD 91,000 - 137,000
Software Engineer
Software Engineer

thehartford • Hartford (CT)

Hybrid
USD 88,000 - 132,000
Staff Data Reliability Engineer - Hybrid
Staff Data Reliability Engineer - Hybrid

The Hartford • Northern (KY)

Hybrid
USD 128,000 - 191,000
Sr. Infra Ops - IT Engineer
Sr. Infra Ops - IT Engineer

thehartford • Hartford (CT)

Hybrid
USD 117,000 - 175,000
Sr. Infra Ops – IT Engineer
Sr. Infra Ops – IT Engineer

The Hartford • Charlotte (NC)

Hybrid
USD 117,000 - 175,000
Staff Software Engineer
Staff Software Engineer

thehartford • Hartford (CT)

Hybrid
USD 116,000 - 174,000