Site Reliability Engineer

Evlo AI

Phoenix (AZ)

On-site

USD 120,000 - 180,000

Full time

5 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to own uptime, performance, and scalability of our production systems. You will shape deployment, monitoring, and scaling across Kubernetes and cloud infra, working with developers to keep services online and engineers shipping safely.

You will automate toil with Python, Bash, and Go, design blameless postmortems, and lead incident response with on-call incident command. A strong foundation in IaC, observability, and reliability is essential.

Qualifications

  • 3–7 years of experience in SRE/DevOps or infra engineering with on-call responsibilities.
  • Hands-on Kubernetes production experience: managing clusters and workloads.
  • IaC proficiency with Terraform and experience with modern CI/CD pipelines.

Responsibilities

  • Build and maintain scalable infrastructure on AWS/GCP using Terraform and code-driven provisioning.
  • Own SLOs and error budgets; define SLIs and build dashboards to drive reliability decisions.
  • Design and operate Kubernetes clusters with autoscaling, POD disruption budgets, and Helm-based deployments.
  • Lead incident response as on-call escalation point; run blameless postmortems and drive remediation.
  • Automate toil with Python, Bash, and Go; implement runbooks and cost-optimization tooling.
  • Harden production systems with network policies, secrets management, and least-privilege IAM.
  • Collaborate with development teams on production readiness and observability coverage.

Skills

On-call experience
Incident response
Problem-solving

Education

BS in Computer Science

Tools

Kubernetes
Terraform
CI/CD (GitHub Actions/GitLab/ArgoCD)
Prometheus/Grafana/Datadog
Python/Bash/Go

Job description

About The Role

The Site Reliability Engineer owns the uptime, performance, and scalability of production systems handling thousands of requests per second. The role blends software engineering with systems thinking - eliminating toil through automation, designing infrastructure that fails gracefully, and running incident response that keeps customers informed and services recovering fast. This is a hands‑on role on a platform team that other engineering teams depend on. You will shape how services are deployed, monitored, and scaled across Kubernetes and cloud infrastructure, and your work directly determines whether engineers can ship safely and customers stay online.


Key Responsibilities


  • Build and maintain scalable infrastructure on AWS/GCP using Terraform, treating everything as code with peer‑reviewed modules and CI‑driven provisioning

  • Own SLOs and error budgets for critical services - define SLIs, build dashboards in Grafana/Datadog, and drive reliability decisions across product teams

  • Design and operate Kubernetes clusters: autoscaling, resource tuning, pod disruption budgets, and Helm‑based service deployment standards

  • Lead incident response as an on‑call escalation point - run incident command, write blameless postmortems, and drive remediation work to closure

  • Automate toil away with Python, Bash, and Go - self‑healing runbooks, capacity management, and cost optimization tooling

  • Harden production systems: implement network policies, secrets management, least‑privilege IAM, and vulnerability remediation pipelines

  • Partner with development teams on production readiness - review designs for reliability, set deployment standards, and improve observability coverage


What We Are Looking For


  • 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including operating large‑scale production systems with real on‑call responsibilities

  • Deep hands‑on experience with Kubernetes in production - operating clusters, debugging workloads, and tuning for performance and cost

  • Strong infrastructure‑as‑code skills with Terraform (or similar) and experience with CI/CD pipelines (GitHub Actions, GitLab CI, or ArgoCD)

  • Solid grasp of distributed systems fundamentals: load balancing, service meshes, queues, caching strategies, and failure modes

  • Production experience with observability stacks: Prometheus, Grafana, Datadog, or equivalent - including building actionable alerts, not noisy ones

  • BS in Computer Science or equivalent practical experience; scripting proficiency in Python, Go, or Bash

  • Bonus: experience with chaos engineering, multi‑region architectures, service mesh (Istio/Linkerd), or database reliability (PostgreSQL, MySQL)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Hidden Jobs • United States

On-site
USD 150,000 - 190,000
Healthcare
Retirement matching
Paid family leave
+3
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Talentify • Houston (TX)

On-site
USD 140,000 - 180,000