Get more replies from employers
Send a job-specific resume in minutes.
Evlo AI is seeking a Site Reliability Engineer to own the reliability, scalability, and performance of production distributed systems across multi-region cloud environments. You will work with software squads to embed resilience, automate toil, and maintain strict SLAs.
You will design and maintain infrastructure on AWS or GCP using Terraform and Pulumi, implement observability with Prometheus, Grafana, OpenTelemetry, and Datadog, and drive incident response with blameless post-mortems.
The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.
The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.
The team works closely with software engineering squads to embed resilience into architecture, automate operational toil, and maintain strict SLAs.