Site Reliability Engineer: Scale & Automation

Evlo AI

San Francisco (CA)

On-site

USD 140,000 - 180,000

Full time

23 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to own production systems, spanning Kubernetes, cloud networking, observability, and automated delivery at scale. You will partner with software and platform teams to reduce toil and improve uptime, latency, and recovery.

You will design highly available infrastructure across AWS/GCP/Azure, build observability with Prometheus and Grafana, and lead incident response with clear post-incident actions.

Qualifications

  • 3–8 years of experience in site reliability engineering or a related role.
  • Hands-on experience operating Kubernetes and containerized workloads in production.
  • Experience with at least one cloud platform (AWS, GCP, or Azure) and infrastructure as code (Terraform).
  • Experience building observability with metrics, logs, and traces (Prometheus, Grafana, OpenTelemetry, Datadog).
  • Programming or scripting ability in Go, Python, or Bash for automation.
  • Incident management, on-call, and root-cause analysis experience.
  • BS in CS/engineering or equivalent; bonus for service meshes, Argo CD, Kafka, multi-region, chaos engineering.

Responsibilities

  • Design and operate highly available production infrastructure on AWS, GCP, or Azure using Kubernetes, Terraform, and managed data services.
  • Build and maintain observability systems to track service health, latency, capacity, and error budgets.
  • Automate deployment, scaling, and recovery workflows through CI/CD pipelines and IaC.
  • Lead incident response for production outages and document post-incident analyses.
  • Define and enforce SLOs, SLIs, and reliability practices across critical services.
  • Improve system performance and resilience with load testing and disaster-recovery exercises.
  • Collaborate on architecture reviews and safe rollout strategies including canary and blue-green deployments.

Skills

Kubernetes operations
Cloud platforms
Automation scripting

Education

Bachelor's degree in CS/engineering or related field

Tools

Terraform
Prometheus
Grafana
OpenTelemetry
Datadog

Job description

Evlo AI is seeking a Site Reliability Engineer to own production systems, spanning Kubernetes, cloud networking, observability, and automated delivery at scale. You will partner with software and platform teams to reduce toil and improve uptime, latency, and recovery.

You will design highly available infrastructure across AWS/GCP/Azure, build observability with Prometheus and Grafana, and lead incident response with clear post-incident actions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Scale an AI‑Powered SaaS Platform
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
DevOps Engineer: AWS, Kubernetes, CI/CD & Reliability
DevOps Engineer: AWS, Kubernetes, CI/CD & Reliability

Evlo AI • Miami (FL)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Cloud SRE: Build Resilient, Secure, Scalable Infra
Cloud SRE: Build Resilient, Secure, Scalable Infra

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Staff Site Reliability Engineer - Scale Global AWS Infra
Staff Site Reliability Engineer - Scale Global AWS Infra

Pearl Street Technologies • Pittsburgh

On-site
USD 130,000 - 180,000
Site Reliability Engineer: AI-Driven Kubernetes Automation
Site Reliability Engineer: AI-Driven Kubernetes Automation

Kindredventures • United States

Remote
USD 8,000 - 15,000
Senior DevOps Engineer: Cloud Infra, Kubernetes & CI/CD
Senior DevOps Engineer: Cloud Infra, Kubernetes & CI/CD

Evlo AI • San Diego (CA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • San Francisco (CA)

On-site
USD 140,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000