Site Reliability Engineer

Unum

Atlanta (GA)

On-site

USD 152,131 - 162,131

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare benefits
401(k) retirement plan with employer 5
Disability insurance
Paid time off
Performance-based incentive plans

Job summary

Unum Group seeks Site Reliability Engineers in Atlanta, GA to design, build, and maintain observability across consumer and client-facing platforms. You will develop dashboards tracking availability, latency, error rate, throughput, capacity, MTTR, and other reliability metrics, while diagnosing distributed systems issues across cloud-based and on‑prem services.

Work with cross‑functional teams to improve service reliability, automate health checks, and mature incident response.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related field with 5 years of experience.
  • 5 years' experience with observability/monitoring platforms (Dynatrace, CloudWatch, Datadog, Grafana, Amplitude).
  • 5 years with containerized and cloud-native architectures in AWS.
  • 5 years supporting incident response, on-call rotations, SLOs/SLAs.
  • 3 years with IaC tools Terraform, CloudFormation, or Ansible.
  • 4 years designing/logging/metrics pipelines; CI/CD with GitHub Actions/Jenkins/Azure DevOps.

Responsibilities

  • Design, build, and maintain observability, monitoring, and alerting capabilities across platforms.
  • Develop dashboards measuring availability, latency, error rate, throughput, and MTTR metrics.
  • Diagnose and troubleshoot distributed systems across cloud-based and on-prem services.
  • Partner with engineering teams to reduce toil and mature incident response practices.
  • Automate service health checks, performance monitoring, and remediation.
  • Manage CI/CD pipeline reliability and deployment quality controls.
  • Conduct root-cause analysis and drive long-term corrective actions.
  • Collaborate with Run teams to productionize monitoring and insights.

Skills

Observability
CloudWatch
Datadog
Grafana
Dynatrace
Amplitude
Python
Bash
PowerShell
Git

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

GitHub
GitLab
Bitbucket
Jenkins
Azure DevOps
Terraform
AWS CloudFormation
Ansible

Job description

Overview

Unum Group seeks Site Reliability Engineers in Atlanta, GA.

Responsibilities
  • Design, build, and maintain observability, monitoring, and alerting capabilities across consumer and client-facing digital platforms.
  • Develop and maintain dashboards that measure availability, latency, error rate, throughput, capacity, MTTR/MBTI, and other reliability metrics.
  • Diagnose and troubleshoot distributed systems issues across cloud-based and on-prem services.
  • Partner with engineering teams to improve service reliability, reduce operational toil, and mature incident response practices.
  • Implement automation for service health checks, performance monitoring, and remediation.
  • Manage CI/CD pipeline reliability and deployment quality controls.
  • Conduct root-cause analysis and drive long-term corrective actions.
  • Collaborate with Run teams to transition monitoring, dashboards, and operational insights into production support processes.
  • Provide guidance on service re-platforming, performance improvements, and architectural decisions based on reliability data.
Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or related field plus 5 years of experience.
  • 5 years of experience with:
    • Observability and monitoring platforms used to monitor application performance and system health using Dynatrace, AWS CloudWatch, Datadog, Grafana, or Amplitude.
    • Working with containerized and cloud-native architectures, including deployment, configuration, and operational support in cloud environments, using AWS.
    • Supporting incident response processes, including participation in on-call rotations, post-incident reviews, and implementation of service-level objectives (SLOs), service-level indicators (SLIs), or service-level agreements (SLAs).
    • Developing scripts or automation to improve system reliability or operational efficiency using Python, Bash, or PowerShell.
    • Troubleshooting distributed systems and analyzing performance bottlenecks across multi-tier or microservices-based architectures.
    • Collaborating with cross-functional engineering teams, including software engineering, platform, infrastructure, or operations teams, within a DevOps or reliability-focused environment.
    • Working with version control systems and collaborative development workflows using GitHub, GitLab, or Bitbucket.
  • 4 years of experience with:
    • Designing, implementing, or maintaining logging, metrics, and distributed tracing pipelines for enterprise or cloud-based systems.
    • Hands‑on experience with continuous integration and continuous deployment (CI/CD) tools and pipelines, using GitHub Actions, Jenkins, or Azure DevOps.
  • 3 years of experience with infrastructure-as-code or configuration management tools to provision, manage, or maintain environments, including Terraform, AWS CloudFormation, or Ansible.
Work Arrangement & Compensation
  • Telecommuting within worksite; up to 5% domestic travel.
  • 40 hours per week.
  • Annual salary range: $152,131 - $162,131 per year (base range $98,340.00 - $201,900.00).
Benefits
  • Healthcare benefits (health, vision, dental), insurance benefits (short & long-term disability), performance-based incentive plans, paid time off, and a 401(k) retirement plan with an employer match up to 5% and an additional 4.5% contribution whether you contribute to the plan or not.
Equal Opportunity Employer

Unum is an equal opportunity employer, considering all qualified applicants and employees for hiring, placement, and advancement, without regard to a person's race, color, religion, national origin, age, genetic information, military status, gender, sexual orientation, gender identity or expression, disability, or protected veteran status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Unum • Dunwoody (GA)

Hybrid
USD 98,000 - 202,000
Site Reliability Engineer: Cloud Observability & Automation
Site Reliability Engineer: Cloud Observability & Automation

Unum • Dunwoody (GA)

Hybrid
USD 98,000 - 202,000
Senior Site Reliability Engineer - Observability & Cloud
Senior Site Reliability Engineer - Observability & Cloud

Unum • Atlanta (GA)

On-site
USD 152,000 - 163,000
Healthcare benefits
401(k) retirement plan with employer 5
Disability insurance
+2
IT Director, Service Stability Operations - Hybrid
IT Director, Service Stability Operations - Hybrid

Unum Insurance • Portland (ME)

Hybrid
USD 109,000 - 224,000
Onsite fitness facilities
Generous paid time off
Employee professional development programs
+2
Principal Data Engineer - CX Analytics
Principal Data Engineer - CX Analytics

Unum • Atlanta (GA)

On-site
USD 109,000 - 224,000
Award-winning culture
Competitive benefits package
Generous PTO
Director - Strategic CX Designer
Director - Strategic CX Designer

Unum • Chattanooga (TN)

On-site
USD 109,000 - 224,000
Healthcare benefits
401(k) retirement plan with employer-m
Paid time off
+1
Data Scientist II - CX Analytics
Data Scientist II - CX Analytics

Unum • Portland (ME)

Hybrid
USD 73,300 - 150,500
Health insurance
Vision insurance
Dental insurance
+5
Data Scientist II - CX Analytics
Data Scientist II - CX Analytics

Unum • Dunwoody (GA)

On-site
USD 73,000 - 151,000
Healthcare benefits
Vision and Dental
Short & Long-Term Disability
+1
Principal Data Scientist - CX Analytics
Principal Data Scientist - CX Analytics

Unum • Portland (ME)

Hybrid
USD 109,100 - 224,000
Health benefits
Vision & Dental insurance
Disability coverage
+2
VP, Platform Services
VP, Platform Services

Unum • Portland (TN)

On-site
USD 202,000 - 416,000
Healthcare benefits
401(k) retirement plan with employer match
Paid time off
+2