Site Reliability Engineer

Alvaria

Town of Texas, Northern (WI, KY)

Hybrid

USD 120,000 - 180,000

Full time

29 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Aspect Software is seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. You will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.

This role partners with application engineering, platform, security, and product teams to define service-level objectives, automate repetitive work, strengthen incident response, and design

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related field.
  • 4–5+ years of experience in site reliability engineering, DevOps, platform engineering, cloud operations, or production software engineering.
  • Hands-on experience operating production workloads in AWS; comparable experience with Azure or GCP is also valuable.
  • Experience with infrastructure as code and configuration automation, such as AWS CDK, CloudFormation, Terraform, or similar tools.
  • Experience building or supporting CI/CD pipelines and source-control workflows, preferably with GitHub and GitHub Actions.
  • Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed-systems failure modes.
  • Experience with observability and monitoring tools such as Datadog, Grafana, CloudWatch or equivalent platforms.
  • Understanding of incident management, root-cause analysis, and blameless post-incident practices.
  • Strong troubleshooting, documentation, collaboration, and communication skills

Responsibilities

  • Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
  • Partner with engineering teams to define and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and actionable alerts
  • Build automation and self-service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
  • Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
  • Participate in an on-call rotation and lead or support incident response, communication, mitigation, and recovery
  • Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes
  • Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
  • Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
  • Strengthen resilience through redundancy, autoscaling, fault-tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
  • Troubleshoot complex issues across applications, infrastructure, networking, databases, and third-party dependencies
  • Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
  • Collaborate across teams and time zones while clearly communicating risk, trade-offs, incident status, and reliability priorities

Skills

Cloud basics
Observability
CI/CD
Incident response
Linux
Networking

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
CloudFormation
GitHub Actions
Datadog
Grafana
CloudWatch
Kubernetes

Job description

If you are unable to complete this application due to a disability, contact this employer to ask for an accommodation or an alternative application process.

Site Reliability Engineer

Indv. Contributor TX, US

2 days ago Requisition ID: 14952

About Aspect Software

Building on more than 50 years of industry experience, Aspect Software is reimagining workforce management through cloud technology, AI, automation, and human‑centered innovation. Our Workforce Engagement Management solutions help organizations solve complex workforce challenges, improve operational performance, and deliver better employee and customer experiences.

We foster a collaborative environment where engineers work across teams and geographies to build secure, scalable, and dependable software. Join us as we modernize intelligent workforce systems used by organizations around the world.

Position Overview

We are seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. In this role, you will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.

This person will partner closely with application engineering, platform, security, and product teams to define service‑level objectives, automate repetitive work, strengthen incident response, and design systems that remain resilient as they scale. The ideal candidate is a pragmatic problem‑solver who measures what matters, learns from failures, and improves systems through automation rather than manual intervention.

Key Responsibilities

  • Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
  • Partner with engineering teams to define and maintain service‑level indicators (SLIs), service‑level objectives (SLOs), error budgets, and actionable alerts
  • Build automation and self‑service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
  • Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
  • Participate in an on‑call rotation and lead or support incident response, communication, mitigation, and recovery
  • Facilitate blameless post‑incident reviews and ensure corrective actions address systemic causes
  • Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
  • Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
  • Strengthen resilience through redundancy, autoscaling, fault‑tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
  • Troubleshoot complex issues across applications, infrastructure, networking, databases, and third‑party dependencies
  • Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
  • Collaborate across teams and time zones while clearly communicating risk, trade‑offs, incident status, and reliability priorities

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 4-5+ years of experience in site reliability engineering, DevOps, platform engineering, cloud operations, or production software engineering
  • Hands‑on experience operating production workloads in AWS; comparable experience with Azure or GCP is also valuable
  • Experience with infrastructure as code and configuration automation, such as AWS CDK, CloudFormation, Terraform, or similar tools
  • Experience building or supporting CI/CD pipelines and source‑control workflows, preferably with GitHub and GitHub Actions
  • Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed‑systems failure modes
  • Experience with observability and monitoring tools such as Datadog, Grafana, CloudWatch or equivalent platforms
  • Understanding of incident management, root‑cause analysis, and blameless post‑incident practices
  • Strong troubleshooting, documentation, collaboration, and communication skills

Preferred Qualifications

  • Experience with containerized or serverless architecture, including Kubernetes, Docker, AWS Lambda, or related technologies
  • Experience with AWS services such as API Gateway, CloudFront, S3, IAM, WAF, DynamoDB, Aurora, EventBridge, SQS, Kinesis, Cognito, VPC, Route 53, and Secrets Manager
  • Experience designing multi‑region systems, disaster recovery strategies, backup and restoration processes, and business‑continuity controls
  • Experience with performance testing, capacity planning, chaos engineering, or reliability testing
  • Familiarity with security, privacy, audit, and compliance requirements for enterprise SaaS products
  • Experience supporting data‑intensive, real‑time, or high‑volume enterprise applications
  • Experience mentoring engineers or leading cross‑team reliability initiatives

Why Join Us?

  • Help shape the reliability practices behind modern workforce technology
  • Work on meaningful, technically challenging systems used by enterprise customers
  • Collaborate with skilled colleagues across engineering, product, security, and operations
  • Influence architecture, tooling, and operational standards as our cloud platforms evolve
  • Access professional development and career‑growth opportunities
  • Join a team that values innovation, accountability, inclusion, and measurable customer outcomes

This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ADP • Town of Texas (WI)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

ALVARIA • Houston (TX)

On-site
USD 120,000 - 160,000
Senior DevOps Engineer
Senior DevOps Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 150,000
DevOps Engineer
DevOps Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

Alvaria Inc • Texas

On-site
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

ADP • Town of Texas (WI)

On-site
USD 90,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

ALVARIA • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive compensation
Career growth
Global collaboration
+1