Site Reliability Engineer

Evlo AI

San Francisco (CA)

On-site

USD 140,000 - 180,000

Full time

15 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to own production systems, spanning Kubernetes, cloud networking, observability, and automated delivery at scale. You will partner with software and platform teams to reduce toil and improve uptime, latency, and recovery.

You will design highly available infrastructure across AWS/GCP/Azure, build observability with Prometheus and Grafana, and lead incident response with clear post-incident actions.

Qualifications

  • 3–8 years of experience in site reliability engineering or a related role.
  • Hands-on experience operating Kubernetes and containerized workloads in production.
  • Experience with at least one cloud platform (AWS, GCP, or Azure) and infrastructure as code (Terraform).
  • Experience building observability with metrics, logs, and traces (Prometheus, Grafana, OpenTelemetry, Datadog).
  • Programming or scripting ability in Go, Python, or Bash for automation.
  • Incident management, on-call, and root-cause analysis experience.
  • BS in CS/engineering or equivalent; bonus for service meshes, Argo CD, Kafka, multi-region, chaos engineering.

Responsibilities

  • Design and operate highly available production infrastructure on AWS, GCP, or Azure using Kubernetes, Terraform, and managed data services.
  • Build and maintain observability systems to track service health, latency, capacity, and error budgets.
  • Automate deployment, scaling, and recovery workflows through CI/CD pipelines and IaC.
  • Lead incident response for production outages and document post-incident analyses.
  • Define and enforce SLOs, SLIs, and reliability practices across critical services.
  • Improve system performance and resilience with load testing and disaster-recovery exercises.
  • Collaborate on architecture reviews and safe rollout strategies including canary and blue-green deployments.

Skills

Kubernetes operations
Cloud platforms
Automation scripting

Education

Bachelor's degree in CS/engineering or related field

Tools

Terraform
Prometheus
Grafana
OpenTelemetry
Datadog

Job description

About The Role

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and predictable at scale. The role spans Kubernetes-based infrastructure, cloud networking, observability, incident response, and automated delivery across distributed systems.

About The Role

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and predictable at scale. The role spans Kubernetes-based infrastructure, cloud networking, observability, incident response, and automated delivery across distributed systems.

You will partner with software engineers and platform teams to reduce operational risk, improve deployment safety, and eliminate recurring sources of toil. The work directly impacts uptime, latency, recovery time, and the ability to ship changes without compromising customer experience.

Key Responsibilities
  • Design and operate highly available production infrastructure on AWS, GCP, or Azure using Kubernetes, Terraform, and managed data services
  • Build and maintain observability systems with Prometheus, Grafana, OpenTelemetry, and centralized logging to track service health, latency, capacity, and error budgets
  • Automate deployment, scaling, and recovery workflows through CI/CD pipelines, infrastructure as code, and self-service platform tooling
  • Lead incident response for production outages, coordinate mitigation, and document clear post-incident analyses with measurable follow-up actions
  • Define and enforce SLOs, SLIs, alerting standards, and reliability practices across critical services
  • Improve system performance and resilience through load testing, capacity planning, failure-mode analysis, and controlled disaster-recovery exercises
  • Collaborate with application engineers on architecture reviews, operational readiness, and safe rollout strategies including canary and blue-green deployments
What We Are Looking For
  • 3–8 years of experience in site reliability engineering, DevOps, production engineering, or a closely related infrastructure role
  • Hands-on experience operating Kubernetes and containerized workloads in production, including cluster security, networking, upgrades, and capacity management
  • Strong proficiency with at least one cloud platform, preferably AWS, GCP, or Azure, and infrastructure as code using Terraform or an equivalent tool
  • Experience building production observability with metrics, logs, and traces using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent
  • Solid programming or scripting ability in Go, Python, or Bash, with a focus on automation, testing, and maintainable operational tooling
  • Practical experience with incident management, SLOs, error budgets, on-call operations, and root-cause analysis for distributed systems
  • Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experience; Bonus: experience with service meshes, Argo CD, Kafka, multi-region architectures, chaos engineering, or compliance-focused infrastructure
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

PrincePerelson and Associates • Salt Lake City (UT)

Hybrid
USD 120,000 - 190,000
Hybrid work schedule (4 days in-office
Medical, dental, vision
Retirement plan
+3
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000