Site Reliability Engineer

Evlo AI

Phoenix (AZ)

On-site

USD 120,000 - 160,000

Full time

5 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Evlo AI is hiring a Site Reliability Engineer to own the reliability and operational maturity of production services on AWS. The role spans Kubernetes, Terraform, CI/CD, observability, incident response, and automation for running distributed systems at scale.

You will help build dependable platform capabilities, define service-level objectives, and orchestrate blameless post-incident reviews to drive measurable improvements.

Qualifications

  • 3–8 years of experience in site reliability engineering or similar roles.
  • Hands-on in AWS with Kubernetes networking, IAM, autoscaling, storage, and prod troubleshooting.
  • Strong infra as code with Terraform and Git-based workflows.
  • Experience building observability and alerting with common tools.
  • Proficiency in Python, Go, or Bash for automation.
  • Bachelor’s degree in CS/engineering or equivalent experience.
  • Nice to have: AWS certs, service mesh, GitOps, chaos engineering, or distributed DB skills.

Responsibilities

  • Design and operate highly available AWS infrastructure with Kubernetes and IaC.
  • Define and enforce service-level objectives and reliability standards.
  • Build and maintain observability and centralized logging platforms.
  • Automate deployment, scaling, backup, recovery, and routine workflows.
  • Lead incident response and post-incident reviews with follow-up actions.
  • Improve CI/CD pipelines with safe deployment controls and testing.
  • Collaborate with engineers to reduce toil and strengthen failure-mode testing.

Skills

Kubernetes
Terraform
CI/CD
Observability
Python
Go
Bash
Git
Linux fundamentals
Networking fundamentals

Education

Bachelor’s degree in CS/Engineering or related field

Tools

Prometheus
Grafana
OpenTelemetry
Datadog

Job description

About The Role

The Site Reliability Engineer owns the reliability, availability, and operational maturity of production services running across AWS. The role spans Kubernetes, Terraform, CI/CD, observability, incident response, and the automation required to operate distributed systems at scale.

About The Role

The Site Reliability Engineer owns the reliability, availability, and operational maturity of production services running across AWS. The role spans Kubernetes, Terraform, CI/CD, observability, incident response, and the automation required to operate distributed systems at scale.

The team is building dependable platform capabilities for engineering teams that ship frequently and serve demanding workloads. This role matters because it turns operational risk into measurable engineering improvements through resilient architecture, clear service-level objectives, and disciplined automation.

Key Responsibilities
  • Design and operate highly available AWS infrastructure using Kubernetes, Terraform, Helm, and infrastructure-as-code best practices
  • Define and enforce service-level objectives, error budgets, and reliability standards across production services
  • Build and maintain observability systems using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms
  • Automate deployment, scaling, backup, recovery, and routine operational workflows through Python, Go, or Bash
  • Lead incident response, coordinate remediation during high-severity events, and produce blameless post-incident reviews with actionable follow-up work
  • Improve CI/CD pipelines with deployment safety controls, automated testing, progressive delivery, and reliable rollback procedures
  • Partner with software engineers to identify performance bottlenecks, eliminate recurring operational toil, and strengthen failure-mode testing
What We Are Looking For
  • 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure roles
  • Hands‑on experience operating Kubernetes workloads in AWS, including networking, IAM, autoscaling, storage, and production troubleshooting
  • Strong proficiency with Terraform or an equivalent infrastructure-as-code framework, plus practical experience managing cloud resources through version control
  • Experience building observability and alerting systems with tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent platforms
  • Proficiency in Python, Go, or Bash for operational automation, tooling, and service integration; familiarity with Linux internals and networking fundamentals
  • Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experience
  • Bonus: Experience with AWS certifications, service mesh technologies, GitOps, chaos engineering, distributed databases, or compliance requirements for production infrastructure
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • San Francisco (CA)

On-site
USD 140,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
DevOps Engineer
DevOps Engineer

Evlo AI • San Diego (CA)

On-site
USD 120,000 - 180,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
DevOps Engineer
DevOps Engineer

Evlo AI • Miami (FL)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000