Site Reliability Engineer

Evlo AI

Denver (CO)

On-site

USD 120,000 - 180,000

Full time

7 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Evlo AI is seeking an experienced Site Reliability/DevOps Engineer to keep large-scale production systems reliable, observable, and fast. You will own Kubernetes clusters, CI/CD pipelines, and the observability stack used by hundreds of engineers to ship daily.

You will define SLOs, instrument everything, automate toil away, and lead blameless postmortems. The role places you on‑call for production services handling significant traffic.

Qualifications

  • 3–7 years of experience in SRE, DevOps, or infrastructure engineering with on‑call rotations.
  • Production Kubernetes ownership: diagnosing and resolving outages and performance issues.
  • Strong infrastructure‑as‑code skills (Terraform) and scripting in Python, Go, or Bash.

Responsibilities

  • Operate and scale Kubernetes-based infrastructure across environments, including upgrades and autoscaling.
  • Build and maintain Terraform modules for provisioning cloud infrastructure (AWS or GCP).
  • Design SLOs, SLIs, and error budgets with product teams and drive reliability improvements.

Skills

SRE/DevOps experience
Kubernetes production experience
Infrastructure as code
Scripting in Python/Go/Bash
Observability tools
Linux networking basics
Incident response & postmortems
GitOps / service meshes (bonus)

Education

BS in Computer Science or equivalent

Tools

Terraform
Kubernetes
GitHub Actions
ArgoCD
Prometheus
Grafana
OpenTelemetry

Job description

About The Role

The role focuses on keeping large-scale production systems reliable, observable, and fast. This team owns the infrastructure layer - Kubernetes clusters, CI/CD pipelines, and the observability stack that hundreds of internal engineers depend on to ship daily.

About The Role

The role focuses on keeping large-scale production systems reliable, observable, and fast. This team owns the infrastructure layer - Kubernetes clusters, CI/CD pipelines, and the observability stack that hundreds of internal engineers depend on to ship daily.

This is a hands‑on position for someone who treats reliability as a product problem: defining SLOs, instrumenting everything, automating away toil, and taking pages when systems misbehave. The team sits directly on‑call for services handling significant production traffic.

Key Responsibilities
  • Operate and scale Kubernetes‑based infrastructure across multiple environments, including cluster upgrades, autoscaling policies, and node lifecycle management
  • Build and maintain Terraform modules for provisioning cloud infrastructure (AWS or GCP), enforcing infrastructure‑as‑code practices across the org
  • Design SLOs, SLIs, and error budgets with product teams, and drive reliability improvements through error budget policy decisions
  • Instrument services with Prometheus, Grafana, and OpenTelemetry; build dashboards and alerts that page on symptoms, not causes
  • Lead incident response for production outages, write blameless postmortems, and follow through on action items to closure
  • Automate operational toil out of existence - runbooks, self‑service tooling, and pipelines that remove manual steps from deploys and scaling operations
  • Improve CI/CD pipelines (GitHub Actions, ArgoCD) to make deployments faster, safer, and progressively rolled out by default
What We Are Looking For
  • 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including ownership of production systems with real on‑call rotations
  • Deep hands‑on experience with Kubernetes in production: operating clusters, troubleshooting workloads, and managing controllers/operators
  • Strong infrastructure‑as‑code skills with Terraform (or equivalent), plus scripting proficiency in Python, Go, or Bash
  • Experience running observability tooling in production: Prometheus, Grafana, distributed tracing (OpenTelemetry or similar)
  • Practical understanding of Linux systems, networking fundamentals (DNS, TLS, load balancing), and container internals
  • Track record of leading incident response and writing postmortems that actually change engineering practice
  • Bonus: Experience with service meshes (Istio/Linkerd), chaos engineering, GitOps workflows, or cloud cost optimization; BS in Computer Science or equivalent practical experience
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DevOps Engineer
DevOps Engineer

Evlo AI • Miami (FL)

On-site
USD 120,000 - 160,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GoGuardian • El Segundo (CA)

Hybrid
USD 180,000 - 240,000