Site Reliability Engineer

Unum

Dunwoody (GA)

Hybrid

USD 98,000 - 202,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Unum, a Fortune 500 insurer, seeks Site Reliability Engineers in Atlanta, GA to design and maintain observability across digital platforms. You will build dashboards, diagnose distributed systems, and partner with engineering to improve reliability and automate health checks.

The role requires strong automation skills, cloud experience on AWS, and experience with incident response and SLOs. Telecommuting w/i worksite with travel up to 5% may apply.

Qualifications

  • Bachelor’s degree in CS/Engineering or related field plus 5 years experience.
  • 5 years with observability/monitoring platforms (Dynatrace, AWS CloudWatch, Datadog, Grafana, Amplitude).
  • 5 years with containerized/cloud-native architectures on AWS.
  • 5 years with incident response, on-call rotations, SLOs/SLAs/SLIs.
  • Python/Bash/PowerShell automation.
  • Troubleshooting distributed systems and microservices.
  • Collaboration with cross-functional DevOps teams.
  • 3 years with IaC tools (Terraform, CloudFormation, Ansible).
  • CI/CD experience with GitHub Actions, Jenkins, or Azure DevOps.

Responsibilities

  • Design, build, and maintain observability and alerting across platforms.
  • Develop dashboards measuring availability, latency, error rate, throughput, MTTR/MBTI.
  • Diagnose distributed systems issues in cloud and on-prem environments.
  • Collaborate with engineering to improve reliability and reduce toil.
  • Automate service health checks and remediation tasks.
  • Maintain CI/CD pipeline reliability and deployment quality controls.
  • Perform root-cause analysis and drive long-term fixes.
  • Work with Run teams to productionize monitoring and insights.
  • Provide guidance on re-platforming and architectural decisions.

Skills

Observability platforms
Cloud-native architectures
Incident response
Automation scripting
Distributed systems
DevOps collaboration
Version control

Education

Bachelor's degree in CS/Engineering or related field

Tools

Dynatrace
AWS CloudWatch
Datadog
Grafana
GitHub/GitLab/Bitbucket

Job description

Job Posting End Date: August 14

Our Fortune 500 company is driving a digital transformation and looking for forward-thinking innovators to disrupt how our industry thinks about and uses technology. As one of the world's leading employee benefits providers, we help millions of people gain affordable access to benefits that help them protect their families, their finances and their futures.

Are you an asker of questions, a solver of problems, and a challenger of the status quo? Our mission is to provide a differentiated customer experience and exceed the expectations people have of technology at any company — not just insurers.

We are seeking individuals to join our team of talented IT professionals who share never-ending passion and an unwavering focus on our customer experience. Team members comfortable working in an agile, fast-paced, and delivery-focused environment thrive in our environment where we value an entrepreneurial spirit and those who challenge the status-quo.

Unum is changing, and we’re excited about what’s next. Join us.

General Summary

Unum Group seeks Site Reliability Engineers in Atlanta, GA.

Applicants who are interested in this position may apply at www.jobpostingtoday.com (Ref #66753) for consideration.

  • Design, build, and maintain observability, monitoring, and alerting capabilities across consumer and client-facing digital platforms
  • Develop and maintain dashboards that measure availability, latency, error rate, throughput, capacity, MTTR/MBTI, and other reliability metrics
  • Diagnose and troubleshoot distributed systems issues across cloud-based and on-prem services
  • Partner with engineering teams to improve service reliability, reduce operational toil, and mature incident response practices
  • Implement automation for service health checks, performance monitoring, and remediation
  • Manage CI/CD pipeline reliability and deployment quality controls
  • Conduct root-cause analysis and drive long-term corrective actions
  • Collaborate with Run teams to transition monitoring, dashboards, and operational insights into production support processes
  • Provide guidance on service re-platforming, performance improvements, and architectural decisions based on reliability data

Requires a Bachelor’s degree in Computer Science, Engineering, or related field plus 5 years of experience. Requires 5 years of experience with the following: Observability and monitoring platforms used to monitor application performance and system health using Dynatrace, AWS CloudWatch, Datadog, Grafana, or Amplitude; Working with containerized and cloud-native architectures, including deployment, configuration, and operational support in cloud environments, using AWS; Supporting incident response processes, including participation in on-call rotations, post-incident reviews, and implementation of service-level objectives (SLOs), service-level indicators (SLIs), or service-level agreements (SLAs); Developing scripts or automation to improve system reliability or operational efficiency using Python, Bash, or PowerShell; Troubleshooting distributed systems and analyzing performance bottlenecks across multi-tier or microservices-based architectures; Collaborating with cross-functional engineering teams, including software engineering, platform, infrastructure, or operations teams, within a DevOps or reliability-focused environment; working with version control systems and collaborative development workflows using GitHub, GitLab, or Bitbucket.

Requires 4 years of experience with the following: Designing, implementing, or maintaining logging, metrics, and distributed tracing pipelines for enterprise or cloud-based systems; Hands-on experience with continuous integration and continuous deployment (CI/CD) tools and pipelines, using GitHub Actions, Jenkins, or Azure DevOps.

Requires 3 years of experience with using infrastructure-as-code or configuration management tools to provision, manage, or maintain environments, including Terraform, AWS CloudFormation, or Ansible. Telecommuting w/i worksite. Up to 5% domestic travel.

40 hours/week; $152,131 - $162,131 per year. This wage range supersedes the base salary range listed below, due to the salary range below reflecting a national range.

~IN1

Our company is built on helping individuals and families, and this starts with our employees. We want employees to maintain a positive balance, which is why we provide access to the benefits and resources they need to invest in themselves. From our onsite fitness facilities and generous paid time off to employee professional development programs, we are committed to helping employees live and work their best – both inside and outside the office.

Unum is an equal opportunity employer, considering all qualified applicants and employees for hiring, placement, and advancement, without regard to a person's race, color, religion, national origin, age, genetic information, military status, gender, sexual orientation, gender identity or expression, disability, or protected veteran status.

The base salary range for applicants for this position is listed below. Unless actual salary is indicated above in the job description, actual pay will be based on skill, geographical location and experience.

$98,340.00-$201,900.00

Additionally, Unum offers a portfolio of benefits and rewards that are competitive and comprehensive including healthcare benefits (health, vision, dental), insurance benefits (short & long-term disability), performance-based incentive plans, paid time off, and a 401(k) retirement plan with an employer match up to 5% and an additional 4.5% contribution whether you contribute to the plan or not. All benefits are subject to the terms and conditions of individual Plans.

Company

Unum

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Unum • Atlanta (GA)

On-site
USD 152,000 - 163,000
Healthcare benefits
401(k) retirement plan with employer 5
Disability insurance
+2
IT Director, Service Stability Operations - Hybrid
IT Director, Service Stability Operations - Hybrid

Unum Insurance • Portland (ME)

Hybrid
USD 109,000 - 224,000
Onsite fitness facilities
Generous paid time off
Employee professional development programs
+2
Director, Value Stream Data & Analytics - Products
Director, Value Stream Data & Analytics - Products

Unum • Chattanooga (TN), Northern (KY)

Hybrid
USD 109,000 - 224,000
Health
Vision
Dental
+6
Data Scientist II - CX Analytics
Data Scientist II - CX Analytics

Unum • Dunwoody (GA)

On-site
USD 73,000 - 151,000
Healthcare benefits
Vision and Dental
Short & Long-Term Disability
+1
Director - Strategic CX Designer
Director - Strategic CX Designer

Unum • Chattanooga (TN)

On-site
USD 109,000 - 224,000
Healthcare benefits
401(k) retirement plan with employer-m
Paid time off
+1
VP, Platform Services
VP, Platform Services

Unum • Portland (TN)

On-site
USD 202,000 - 416,000
Healthcare benefits
401(k) retirement plan with employer match
Paid time off
+2
Data Scientist II - CX Analytics
Data Scientist II - CX Analytics

Unum • Portland (ME)

Hybrid
USD 73,300 - 150,500
Health insurance
Vision insurance
Dental insurance
+5
Benefit Specialist Trainee - Chattanooga
Benefit Specialist Trainee - Chattanooga

Unum • Chattanooga (TN)

On-site
USD 40,000 - 76,000
Award-winning culture
Inclusion and diversity as a priority
Performance Based Incentive Plans
+9
CX Manager
CX Manager

Unum • Columbia (SC)

On-site
USD 89,000 - 184,000
Health benefits
Vision benefits
Dental benefits
+8
Principal Data Scientist - CX Analytics
Principal Data Scientist - CX Analytics

Unum • Portland (ME)

Hybrid
USD 109,100 - 224,000
Health benefits
Vision & Dental insurance
Disability coverage
+2