SRE Architect: Build Reliable, Scalable Cloud Infra

Evlo AI

Chicago (IL)

On-site

USD 120,000 - 180,000

Full time

26 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to own the reliability, scalability, and operational readiness of production systems on AWS and Kubernetes. You will automate infrastructure, improve observability, and lead incident response to minimize downtime and downtime impact.

You will collaborate with platform, application, and security engineers to improve service availability, deployment safety, and recovery time.

Qualifications

  • Bachelor's degree in CS/engineering or related field, or equivalent practical experience.

Responsibilities

  • Build and maintain highly available infrastructure on AWS using Terraform, Kubernetes, and automated CI/CD pipelines
  • Define and improve service-level objectives, error budgets, and operational metrics for critical production services
  • Develop observability solutions with Prometheus, Grafana, OpenTelemetry, and centralized logging platforms such as ELK or Datadog
  • Lead incident response, coordinate technical recovery, and produce blameless postmortems with actionable follow-up work
  • Automate provisioning, configuration management, deployments, and routine operational tasks using Python, Go, or Bash
  • Harden production environments through capacity planning, disaster recovery testing, access controls, and infrastructure security practices
  • Partner with software teams to improve system design, release processes, performance, and operational readiness before launch

Skills

SRE/DevOps
Platform engineering
Production infra
Incident response
Observability design
Automation scripting
Infrastructure as Code
Capacity planning
Disaster recovery
Release coordination
Kubernetes ops
Cloud fundamentals

Education

Bachelor’s degree or equivalent

Tools

Terraform
Kubernetes
OpenTelemetry
Prometheus
Grafana
ELK
Datadog
Python
Go
Bash
CI/CD pipelines
Argo CD
Helm
Kafka
PostgreSQL

Job description

Evlo AI is seeking a Site Reliability Engineer to own the reliability, scalability, and operational readiness of production systems on AWS and Kubernetes. You will automate infrastructure, improve observability, and lead incident response to minimize downtime and downtime impact.

You will collaborate with platform, application, and security engineers to improve service availability, deployment safety, and recovery time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud SRE: Build Resilient, Secure, Scalable Infra
Cloud SRE: Build Resilient, Secure, Scalable Infra

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
SRE Architecture Lead: Reliability & Cloud Platform
SRE Architecture Lead: Reliability & Cloud Platform

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Senior SRE Engineer: Build Resilient, Automated Cloud
Senior SRE Engineer: Build Resilient, Automated Cloud

eMed LLC. • Miami (FL)

On-site
USD 140,000 - 190,000
Retirement Plan (401k with Company Map
Life Insurance (Basic, Voluntary & AD
Paid Time Off
+3
Site Reliability Engineer
Site Reliability Engineer

Involved Solutions • Austin (TX)

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer - Scale, Observability & HA
Senior Site Reliability Engineer - Scale, Observability & HA

Involved Solutions • Austin (TX)

On-site
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
SRE: Cloud Reliability & Automation Engineer
SRE: Cloud Reliability & Automation Engineer

Retool, Inc. • San Francisco (CA)

Hybrid
USD 163,000 - 306,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7