Site Reliability Engineer

Evlo AI

Denver (CO)

On-site

USD 120,000 - 180,000

Full time

12 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Evlo AI in Denver is seeking a Site Reliability Engineer to design, automate, and operate scalable infrastructure behind high‑availability services. You will work with Kubernetes, cloud platforms, and deployment pipelines to improve reliability, observability, and performance.

You will partner with developers to reduce toil, define SLOs, and lead incident response with a focus on durable engineering improvements. This hands-on role requires strong debugging skills and collaboration across teams.

Qualifications

  • 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or related infra role.
  • Hands-on experience operating production workloads on AWS, GCP, or Azure.
  • Strong Kubernetes experience including deployments, services, ingress, resource mgmt, troubleshooting, upgrades.
  • Proficiency with infrastructure-as-code tools (Terraform) and version control for cloud environments.
  • Working knowledge of Linux, TCP/IP networking, distributed systems, and Docker.
  • Experience building observability and incident-management practices around metrics, logs, traces, on-call workflows, and post-incident reviews.
  • Bonus: Go, service meshes, Argo CD, chaos engineering, multi-region architectures, degree in CS/engineering.

Responsibilities

  • Design and maintain highly available infrastructure across AWS or GCP using Kubernetes, Terraform, and IaC practices.
  • Build and improve CI/CD pipelines with GitHub Actions, Argo CD, Jenkins, or equivalents.
  • Define and manage service-level objectives, error budgets, alerting standards, and operational metrics for critical services.
  • Implement observability using Prometheus, Grafana, OpenTelemetry, Elasticsearch, or comparable logging/monitoring platforms.
  • Lead incident response, perform structured root-cause analysis, and deliver improvements to reduce recurrence and recovery time.
  • Automate provisioning, deployment, scaling, and routine tasks with Python, Go, or shell scripting.
  • Collaborate on capacity planning, disaster recovery, performance testing, and production readiness reviews.

Skills

Kubernetes
Cloud platforms
Terraform
CI/CD pipelines
Linux
Networking
Observability
Incident response
Python/Go scripting
On-call

Education

Degree in Computer Science or Engineering

Tools

Terraform
GitHub Actions
Argo CD
Jenkins
Docker

Job description

About The Role

The Site Reliability Engineer will design, automate, and operate the infrastructure behind highly available production services. The role spans Kubernetes, cloud platforms, observability, incident response, and deployment systems, with a focus on making distributed systems more reliable, measurable, and easier to operate at scale.

About The Role

The Site Reliability Engineer will design, automate, and operate the infrastructure behind highly available production services. The role spans Kubernetes, cloud platforms, observability, incident response, and deployment systems, with a focus on making distributed systems more reliable, measurable, and easier to operate at scale.

The engineer will partner with application developers and platform engineers to eliminate recurring operational work, improve service-level objectives, and build systems that remain resilient under growth and failure. This is a hands‑on role for someone who is comfortable debugging complex production issues and turning those lessons into durable engineering improvements.

Key Responsibilities
  • Design and maintain highly available infrastructure across AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code practices
  • Build and improve CI/CD pipelines with tools such as GitHub Actions, Argo CD, Jenkins, or equivalent systems
  • Define and manage service-level objectives, error budgets, alerting standards, and operational metrics for critical services
  • Implement observability using Prometheus, Grafana, OpenTelemetry, Elasticsearch, or comparable logging and monitoring platforms
  • Lead incident response, perform structured root‑cause analysis, and deliver follow‑up improvements that reduce recurrence and recovery time
  • Automate provisioning, deployment, scaling, and routine operational tasks with Python, Go, or shell scripting
  • Collaborate on capacity planning, disaster recovery, performance testing, and production readiness reviews
What We Are Looking For
  • 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or a closely related infrastructure role
  • Hands‑on experience operating production workloads on AWS, GCP, or Azure, including networking, compute, storage, IAM, and managed services
  • Strong Kubernetes experience, including deployments, services, ingress, resource management, troubleshooting, and upgrades
  • Proficiency with Terraform or a comparable infrastructure‑as‑code tool and practical experience managing cloud environments through version control
  • Working knowledge of Linux systems, TCP/IP networking, distributed systems, and container technologies such as Docker
  • Experience building observability and incident‑management practices around metrics, logs, traces, on‑call workflows, and post‑incident reviews
  • Bonus: Experience with Go, service meshes, Argo CD, chaos engineering, multi‑region architectures, and formal training or a degree in computer science, engineering, or a related field
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Seattle (WA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000