Site Reliability Engineer

Evlo AI

San Francisco (CA)

On-site

USD 140,000 - 200,000

Full time

11 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evlo AI is seeking a Site Reliability Engineer to own the reliability, scalability, and performance of production distributed systems across multi-region cloud environments. You will work with software squads to embed resilience, automate toil, and maintain strict SLAs.

You will design and maintain infrastructure on AWS or GCP using Terraform and Pulumi, implement observability with Prometheus, Grafana, OpenTelemetry, and Datadog, and drive incident response with blameless post-mortems.

Qualifications

  • 3–7 years of experience in SRE/DevOps or systems engineering in high-growth tech environments.
  • Strong Kubernetes administration and containerization expertise.
  • Proficiency in Python or Go for automation and tooling.
  • Experience with CI/CD, GitOps, and IaC frameworks.
  • Understanding of networking and distributed systems failure modes.

Responsibilities

  • Design, build, and maintain production infrastructure on AWS or GCP using IaC (Terraform, Pulumi).
  • Implement observability stacks (Prometheus, Grafana, OpenTelemetry, Datadog) for real-time monitoring and alerts.
  • Lead incident response and blameless postmortems to improve resilience.
  • Automate deployments with GitHub Actions, ArgoCD, and Kubernetes.
  • Optimize cloud costs, resources, and database performance without sacrificing reliability.
  • Enforce security, IAM policies, and disaster recovery runbooks across environments.

Skills

Kubernetes administration
Containerization (Docker)
Python or Go
CI/CD pipelines
Networking fundamentals
Distributed systems

Tools

Terraform
Pulumi
Prometheus
Grafana
OpenTelemetry
Datadog
GitHub Actions
ArgoCD

Job description

About The Role

The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.

About The Role

The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.

The team works closely with software engineering squads to embed resilience into architecture, automate operational toil, and maintain strict SLAs.

Key Responsibilities
  • Design, build, and maintain production infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform and Pulumi
  • Implement comprehensive observability stacks using Prometheus, Grafana, OpenTelemetry, and Datadog for real-time monitoring and alerting
  • Drive incident response and conduct thorough blameless post-mortems to continuously improve system resilience and prevent recurrence
  • Automate deployment pipelines and release engineering processes using GitHub Actions, ArgoCD, and Kubernetes
  • Optimize cloud infrastructure costs, resource utilization, and database performance without compromising system reliability
  • Establish and enforce security compliance, IAM policies, and disaster recovery runbooks across all production environments
What We Are Looking For
  • 3–7 years of experience in Site Reliability Engineering, DevOps, or systems engineering roles within high-growth tech environments
  • Deep expertise in Kubernetes administration, containerization (Docker), and service mesh architectures
  • Strong proficiency in scripting languages such as Python or Go for automation and tool development
  • Hands-on experience with modern CI/CD pipelines, GitOps workflows, and infrastructure as code frameworks
  • Solid understanding of networking fundamentals (TCP/IP, DNS, TLS, load balancing) and distributed systems failure modes
  • Bonus: Experience with Chaos Engineering practices, eBPF for observability, or holding professional cloud certifications (AWS/GCP)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 180,000