Site Reliability Engineer

Evlo AI

Chicago (IL)

On-site

USD 120,000 - 180,000

Full time

18 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to own the reliability, scalability, and operational readiness of production systems on AWS and Kubernetes. You will automate infrastructure, improve observability, and lead incident response to minimize downtime and downtime impact.

You will collaborate with platform, application, and security engineers to improve service availability, deployment safety, and recovery time.

Qualifications

  • Bachelor's degree in CS/engineering or related field, or equivalent practical experience.

Responsibilities

  • Build and maintain highly available infrastructure on AWS using Terraform, Kubernetes, and automated CI/CD pipelines
  • Define and improve service-level objectives, error budgets, and operational metrics for critical production services
  • Develop observability solutions with Prometheus, Grafana, OpenTelemetry, and centralized logging platforms such as ELK or Datadog
  • Lead incident response, coordinate technical recovery, and produce blameless postmortems with actionable follow-up work
  • Automate provisioning, configuration management, deployments, and routine operational tasks using Python, Go, or Bash
  • Harden production environments through capacity planning, disaster recovery testing, access controls, and infrastructure security practices
  • Partner with software teams to improve system design, release processes, performance, and operational readiness before launch

Skills

SRE/DevOps
Platform engineering
Production infra
Incident response
Observability design
Automation scripting
Infrastructure as Code
Capacity planning
Disaster recovery
Release coordination
Kubernetes ops
Cloud fundamentals

Education

Bachelor’s degree or equivalent

Tools

Terraform
Kubernetes
OpenTelemetry
Prometheus
Grafana
ELK
Datadog
Python
Go
Bash
CI/CD pipelines
Argo CD
Helm
Kafka
PostgreSQL

Job description

About The Role

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production systems running across AWS and Kubernetes. The role focuses on infrastructure automation, observability, incident response, and eliminating recurring operational work through engineering.

About The Role

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production systems running across AWS and Kubernetes. The role focuses on infrastructure automation, observability, incident response, and eliminating recurring operational work through engineering.

You will partner with platform, application, and security engineers to improve service availability, deployment safety, and recovery time. The team manages systems where measurable uptime, predictable performance, and disciplined production operations are critical to customer experience.

Key Responsibilities
  • Build and maintain highly available infrastructure on AWS using Terraform, Kubernetes, and automated CI/CD pipelines
  • Define and improve service-level objectives, error budgets, and operational metrics for critical production services
  • Develop observability solutions with Prometheus, Grafana, OpenTelemetry, and centralized logging platforms such as ELK or Datadog
  • Lead incident response, coordinate technical recovery, and produce blameless postmortems with actionable follow-up work
  • Automate provisioning, configuration management, deployments, and routine operational tasks using Python, Go, or Bash
  • Harden production environments through capacity planning, disaster recovery testing, access controls, and infrastructure security practices
  • Partner with software teams to improve system design, release processes, performance, and operational readiness before launch
What We Are Looking For
  • 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure roles
  • Hands-on experience operating Kubernetes workloads and cloud infrastructure, preferably AWS services including EC2, EKS, IAM, RDS, S3, and CloudWatch
  • Strong Infrastructure as Code experience with Terraform, including reusable modules, state management, and automated change workflows
  • Proficiency with Linux systems, networking fundamentals, containers, and scripting in Python, Go, or Bash
  • Experience building production observability and alerting practices using tools such as Prometheus, Grafana, Datadog, OpenTelemetry, or ELK
  • Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent practical experience
  • Bonus: Experience with service mesh technologies, Argo CD, Helm, Kafka, PostgreSQL, chaos engineering, or formal SRE practices
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Devops Engineer
Devops Engineer

GCS Recruitment • Cherry Hill Township (NJ)

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veriipro • Charlotte (NC)

On-site
USD 140,000 - 190,000
Site Reliability Specialist
Site Reliability Specialist

GCS Recruitment • Philadelphia

On-site
USD 110,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000