Site Reliability Engineer: Scale Uptime & Automate Kubernetes

Evlo AI

Phoenix (AZ)

On-site

USD 120,000 - 180,000

Full time

21 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to own uptime, performance, and scalability of our production systems. You will shape deployment, monitoring, and scaling across Kubernetes and cloud infra, working with developers to keep services online and engineers shipping safely.

You will automate toil with Python, Bash, and Go, design blameless postmortems, and lead incident response with on-call incident command. A strong foundation in IaC, observability, and reliability is essential.

Qualifications

  • 3–7 years of experience in SRE/DevOps or infra engineering with on-call responsibilities.
  • Hands-on Kubernetes production experience: managing clusters and workloads.
  • IaC proficiency with Terraform and experience with modern CI/CD pipelines.

Responsibilities

  • Build and maintain scalable infrastructure on AWS/GCP using Terraform and code-driven provisioning.
  • Own SLOs and error budgets; define SLIs and build dashboards to drive reliability decisions.
  • Design and operate Kubernetes clusters with autoscaling, POD disruption budgets, and Helm-based deployments.
  • Lead incident response as on-call escalation point; run blameless postmortems and drive remediation.
  • Automate toil with Python, Bash, and Go; implement runbooks and cost-optimization tooling.
  • Harden production systems with network policies, secrets management, and least-privilege IAM.
  • Collaborate with development teams on production readiness and observability coverage.

Skills

On-call experience
Incident response
Problem-solving

Education

BS in Computer Science

Tools

Kubernetes
Terraform
CI/CD (GitHub Actions/GitLab/ArgoCD)
Prometheus/Grafana/Datadog
Python/Bash/Go

Job description

Evlo AI is seeking a Site Reliability Engineer to own uptime, performance, and scalability of our production systems. You will shape deployment, monitoring, and scaling across Kubernetes and cloud infra, working with developers to keep services online and engineers shipping safely.

You will automate toil with Python, Bash, and Go, design blameless postmortems, and lead incident response with on-call incident command. A strong foundation in IaC, observability, and reliability is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer - Cloud, Kubernetes & CI/CD
Senior DevOps Engineer - Cloud, Kubernetes & CI/CD

Evlo AI • Boston (MA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer: AI-Driven Kubernetes Automation
Site Reliability Engineer: AI-Driven Kubernetes Automation

Kindredventures • United States

Remote
USD 8,504 - 14,882
Senior Site Reliability Engineer – Scalable AI Infra
Senior Site Reliability Engineer – Scalable AI Infra

Tavily Inc. • New York (NY)

Hybrid
USD 156,000 - 262,000
100% company-paid medical, dental, and vision coverage
Up to 4% company match 401(k) plan
20 weeks paid parental leave for primary caregivers
+2
Staff Site Reliability Engineer - Scale Global AWS Infra
Staff Site Reliability Engineer - Scale Global AWS Infra

Pearl Street Technologies • Pittsburgh

On-site
USD 130,000 - 180,000