Site Reliability Engineer

ALVARIA

Houston (TX)

On-site

USD 120,000 - 160,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Aspect Software is seeking a Site Reliability Engineer to design, build, and operate reliable cloud services. You will partner with engineering, platform, security, and product teams to define SLOs, automate toil, and improve incident response.

The role emphasizes pragmatic problem-solving, automation over manual effort, and managing reliability as scale increases. The ideal candidate will drive CI/CD improvements, maintain runbooks, and collaborate across time zones to ensure service resilience

Qualifications

  • Bachelor's degree in CS, Engineering, or related field, or equivalent experience.
  • 4–5+ years in SRE/DevOps, cloud operations, or production software engineering.
  • Hands-on AWS production experience; Azure or GCP valuable.
  • Proficiency in Python, Go, TS/JS or C#.
  • IaC and automation with AWS CDK, CloudFormation, Terraform.
  • CI/CD workflows, GitHub Actions experience.
  • Linux, networking, DNS, load balancing, security basics.
  • Observability tools: Datadog, Grafana, CloudWatch.
  • Incident management and blameless post-incident practices.

Responsibilities

  • Design, build, and operate highly available, scalable production services.
  • Define SLIs/SLOs, error budgets, and actionable alerts.
  • Build automation and self-service tooling to reduce toil.
  • Improve observability with metrics, logs, tracing, dashboards.
  • Participate in on-call rotation and lead incident response.
  • Conduct blameless post-incident reviews and address root causes.
  • Improve CI/CD with validation, deployment gates, canaries, rollback.
  • Review designs for reliability, security, cost efficiency.
  • Enhance resilience with autoscaling, disaster recovery planning.

Skills

Python
Go
TypeScript/JavaScript
C#
Linux
Networking
Troubleshooting
Documentation
Communication

Education

Bachelor's degree in Computer Science or related field

Tools

AWS CDK
CloudFormation
Terraform
GitHub Actions
Datadog
Grafana
CloudWatch
Kubernetes
Docker

Job description

About Aspect Software

Building on more than 50 years of industry experience, Aspect Software is reimagining workforce management through cloud technology, AI, automation, and human-centered innovation. Our Workforce Engagement Management solutions help organizations solve complex workforce challenges, improve operational performance, and deliver better employee and customer experiences.

We foster a collaborative environment where engineers work across teams and geographies to build secure, scalable, and dependable software. Join us as we modernize intelligent workforce systems used by organizations around the world.

Position Overview

We are seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. In this role, you will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.

This person will partner closely with application engineering, platform, security, and product teams to define service-level objectives, automate repetitive work, strengthen incident response, and design systems that remain resilient as they scale. The ideal candidate is a pragmatic problem-solver who measures what matters, learns from failures, and improves systems through automation rather than manual intervention.

Key Responsibilities

  • Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
  • Partner with engineering teams to define and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and actionable alerts
  • Build automation and self-service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
  • Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
  • Participate in an on-call rotation and lead or support incident response, communication, mitigation, and recovery
  • Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes
  • Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
  • Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
  • Strengthen resilience through redundancy, autoscaling, fault-tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
  • Troubleshoot complex issues across applications, infrastructure, networking, databases, and third-party dependencies
  • Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
  • Collaborate across teams and time zones while clearly communicating risk, trade-offs, incident status, and reliability priorities

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 4-5+ years of experience in site reliability engineering, DevOps, platform engineering, cloud operations, or production software engineering
  • Hands-on experience operating production workloads in AWS; comparable experience with Azure or GCP is also valuable
  • Proficiency in at least one programming or scripting language such as Python, Go, TypeScript/JavaScript, or C#
  • Experience with infrastructure as code and configuration automation, such as AWS CDK, CloudFormation, Terraform, or similar tools
  • Experience building or supporting CI/CD pipelines and source-control workflows, preferably with GitHub and GitHub Actions
  • Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed-systems failure modes
  • Experience with observability and monitoring tools such as Datadog, Grafana, CloudWatch or equivalent platforms
  • Understanding of incident management, root-cause analysis, and blameless post-incident practices
  • Strong troubleshooting, documentation, collaboration, and communication skills

Preferred Qualifications

  • Experience with containerized or serverless architecture, including Kubernetes, Docker, AWS Lambda, or related technologies
  • Experience with AWS services such as API Gateway, CloudFront, S3, IAM, WAF, DynamoDB, Aurora, EventBridge, SQS, Kinesis, Cognito, VPC, Route 53, and Secrets Manager
  • Experience designing multi-region systems, disaster recovery strategies, backup and restoration processes, and business-continuity controls
  • Experience with performance testing, capacity planning, chaos engineering, or reliability testing
  • Familiarity with security, privacy, audit, and compliance requirements for enterprise SaaS products
  • Experience supporting data-intensive, real-time, or high-volume enterprise applications
  • Experience mentoring engineers or leading cross-team reliability initiatives

Why Join Us?

  • Help shape the reliability practices behind modern workforce technology
  • Work on meaningful, technically challenging systems used by enterprise customers
  • Collaborate with skilled colleagues across engineering, product, security, and operations
  • Influence architecture, tooling, and operational standards as our cloud platforms evolve
  • Access professional development and career-growth opportunities
  • Join a team that values innovation, accountability, inclusion, and measurable customer outcomes

This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

ADP • Town of Texas (WI)

On-site
USD 120,000 - 150,000
Senior Cloud Reliability Engineer
Senior Cloud Reliability Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

ALVARIA • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive compensation
Career growth
Global collaboration
+1
Cloud Site Reliability Engineer: Build Resilient SaaS
Cloud Site Reliability Engineer: Build Resilient SaaS

ADP • Town of Texas (WI)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
DevOps Engineer
DevOps Engineer

Alvaria Inc • Texas

On-site
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 160,000
Senior DevOps Engineer
Senior DevOps Engineer

Alvaria • Town of Texas (WI), Northern (KY)

Hybrid
USD 120,000 - 150,000
DevOps Engineer
DevOps Engineer

ALVARIA • Houston (TX)

On-site
USD 90,000 - 130,000