Site Reliability Engineer

Evlo AI

Seattle (WA)

On-site

USD 140,000 - 190,000

Full time

12 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Evlo AI is seeking a Site Reliability Engineer to design, operate, and improve production infrastructure on Kubernetes and AWS, with a focus on availability, latency, and safety.

You will own production health, define SLIs/SLOs, and collaborate with software teams to build reliable deployment and recovery workflows, influencing platform architecture and engineering standards.

Qualifications

  • Hands-on experience operating Kubernetes in production.
  • Experience building and maintaining observability systems with metrics, logs, traces, and dashboards.
  • Bachelor's degree in CS/Engineering or equivalent practical experience.

Responsibilities

  • Operate and improve highly available production services on AWS and Kubernetes.
  • Define and maintain SLIs, SLOs, error budgets, and dashboards.
  • Automate provisioning and configuration with Terraform, Helm, and CI/CD pipelines.
  • Lead incident response and post-incident reviews with actionable follow-ups.
  • Harden deployment pipelines with progressive delivery, health checks, and change management.

Skills

Kubernetes operations
Observability
Incident response
Automation

Education

Bachelor's degree in CS/Engineering

Tools

Terraform
Helm
GitHub Actions
Prometheus
Grafana
OpenTelemetry
Datadog

Job description

About The Role

The Site Reliability Engineer will design, operate, and improve the production infrastructure supporting high-volume, customer-facing services. The role focuses on Kubernetes, AWS, observability, incident response, and automation across distributed systems where availability, latency, and operational safety are critical.

About The Role

The Site Reliability Engineer will design, operate, and improve the production infrastructure supporting high-volume, customer-facing services. The role focuses on Kubernetes, AWS, observability, incident response, and automation across distributed systems where availability, latency, and operational safety are critical.
The engineer will partner with software teams to define service-level objectives, eliminate recurring failure modes, and build reliable deployment and recovery workflows. This role has direct ownership of production health and meaningful influence over platform architecture, engineering standards, and operational practices.

Key Responsibilities
  • Operate and improve highly available production services running on AWS and Kubernetes, including capacity planning, scaling, and failure recovery
  • Define and maintain SLIs, SLOs, error budgets, and operational dashboards using tools such as Prometheus, Grafana, and OpenTelemetry
  • Automate infrastructure provisioning and configuration with Terraform, Helm, and GitHub Actions or equivalent CI/CD systems
  • Lead incident response, coordinate technical mitigation, and produce clear post-incident reviews with measurable corrective actions
  • Harden deployment pipelines with progressive delivery, automated rollback, health checks, and change-management controls
  • Identify and eliminate recurring toil through Python or Go automation, platform tooling, and self-service workflows
  • Collaborate with application engineers on performance tuning, resilience testing, dependency management, and production readiness reviews
What We Are Looking For
  • 3-8 years of experience in site reliability engineering, DevOps, platform engineering, or a closely related production infrastructure role
  • Hands-on experience operating Kubernetes workloads in production, including deployments, networking, storage, ingress, and troubleshooting
  • Strong knowledge of AWS services such as EC2, EKS, IAM, VPC, RDS, S3, and CloudWatch
  • Proficiency with infrastructure as code and delivery tooling, including Terraform, Helm, Git, and CI/CD pipelines
  • Experience building observability systems with metrics, logs, traces, alerting, and on-call practices using tools such as Prometheus, Grafana, Datadog, or OpenTelemetry
  • Proficiency in Python, Go, or a similar programming language, with a track record of replacing manual operational work with reliable automation
  • Bachelor's degree in computer science, engineering, or a related technical field, or equivalent practical experience
  • Bonus: experience with service meshes, distributed systems, chaos engineering, compliance-focused infrastructure, or multi-region production environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

Evlo AI • Seattle (WA)

Hybrid
USD 120,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000