Site Reliability Engineer

ADP

Town of Texas (WI)

On-site

USD 120,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Aspect Software is seeking a Site Reliability Engineer to build and operate reliable, scalable cloud services with a focus on observability, automation, and incident management.

You will collaborate with cross-functional teams to define SLOs/SLIs, automate toil, and strengthen incident response across SaaS platforms. The role emphasizes resilient deployment practices and continuous improvement of reliability and performance.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
  • 4-5+ years of experience in site reliability engineering, DevOps, or production software engineering.
  • Hands-on experience operating production workloads in AWS; Azure or GCP is valuable.
  • Experience with infrastructure as code and configuration automation (e.g., AWS CDK, CloudFormation, Terraform).
  • Experience building or supporting CI/CD pipelines and source-control workflows (GitHub).
  • Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed-systems failure modes.
  • Observability and monitoring experience (Datadog, Grafana, CloudWatch).
  • Understanding of incident management, root-cause analysis, and blameless post-incident practices.

Responsibilities

  • Design, build, and operate highly available, scalable production services and cloud infrastructure.
  • Define and maintain SLIs, SLOs, error budgets, and actionable alerts with engineering teams.
  • Develop automation to reduce toil and improve deployment reliability.
  • Improve observability with metrics, logs, tracing, dashboards, and health checks.
  • Participate in on-call rotations and lead incident response and post-incident reviews.
  • Enhance CI/CD pipelines with automated validation, progressive delivery, canary releases, and reliable rollbacks.
  • Review designs for reliability, capacity, security, and cost efficiency.

Tools

AWS
Terraform
GitHub Actions
Datadog
Grafana

Job description

If you are unable to complete this application due to a disability, contact this employer to ask for an accommodation or an alternative application process.

Site Reliability Engineer

Indv. Contributor TX, US

2 days ago Requisition ID: 14952

About Aspect Software

Building on more than 50 years of industry experience, Aspect Software is reimagining workforce management through cloud technology, AI, automation, and human‑centered innovation. Our Workforce Engagement Management solutions help organizations solve complex workforce challenges, improve operational performance, and deliver better employee and customer experiences.

We foster a collaborative environment where engineers work across teams and geographies to build secure, scalable, and dependable software. Join us as we modernize intelligent workforce systems used by organizations around the world.

Position Overview

We are seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. In this role, you will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.

This person will partner closely with application engineering, platform, security, and product teams to define service‑level objectives, automate repetitive work, strengthen incident response, and design systems that remain resilient as they scale. The ideal candidate is a pragmatic problem‑solver who measures what matters, learns from failures, and improves systems through automation rather than manual intervention.

Key Responsibilities

  • Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
  • Partner with engineering teams to define and maintain service‑level indicators (SLIs), service‑level objectives (SLOs), error budgets, and actionable alerts
  • Build automation and self‑service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
  • Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
  • Participate in an on‑call rotation and lead or support incident response, communication, mitigation, and recovery
  • Facilitate blameless post‑incident reviews and ensure corrective actions address systemic causes
  • Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
  • Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
  • Strengthen resilience through redundancy, autoscaling, fault‑tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
  • Troubleshoot complex issues across applications, infrastructure, networking, databases, and third‑party dependencies
  • Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
  • Collaborate across teams and time zones while clearly communicating risk, trade‑offs, incident status, and reliability priorities

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 4-5+ years of experience in site reliability engineering, DevOps, platform engineering, cloud operations, or production software engineering
  • Hands‑on experience operating production workloads in AWS; comparable experience with Azure or GCP is also valuable
  • Experience with infrastructure as code and configuration automation, such as AWS CDK, CloudFormation, Terraform, or similar tools
  • Experience building or supporting CI/CD pipelines and source‑control workflows, preferably with GitHub and GitHub Actions
  • Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed‑systems failure modes
  • Experience with observability and monitoring tools such as Datadog, Grafana, CloudWatch or equivalent platforms
  • Understanding of incident management, root‑cause analysis, and blameless post‑incident practices
  • Strong troubleshooting, documentation, collaboration, and communication skills

Preferred Qualifications

  • Experience with containerized or serverless architecture, including Kubernetes, Docker, AWS Lambda, or related technologies
  • Experience with AWS services such as API Gateway, CloudFront, S3, IAM, WAF, DynamoDB, Aurora, EventBridge, SQS, Kinesis, Cognito, VPC, Route 53, and Secrets Manager
  • Experience designing multi‑region systems, disaster recovery strategies, backup and restoration processes, and business‑continuity controls
  • Experience with performance testing, capacity planning, chaos engineering, or reliability testing
  • Familiarity with security, privacy, audit, and compliance requirements for enterprise SaaS products
  • Experience supporting data‑intensive, real‑time, or high‑volume enterprise applications
  • Experience mentoring engineers or leading cross‑team reliability initiatives

Why Join Us?

  • Help shape the reliability practices behind modern workforce technology
  • Work on meaningful, technically challenging systems used by enterprise customers
  • Collaborate with skilled colleagues across engineering, product, security, and operations
  • Influence architecture, tooling, and operational standards as our cloud platforms evolve
  • Access professional development and career‑growth opportunities
  • Join a team that values innovation, accountability, inclusion, and measurable customer outcomes

This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 150,000
DevOps Engineer
DevOps Engineer

Alvaria Inc • Texas

On-site
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

ADP • Town of Texas (WI)

On-site
USD 90,000 - 140,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Pacificacontinental • Pacifica (CA)

On-site
USD 140,000 - 190,000