An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Clearwater Analytics is seeking a Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of cloud-native systems. You will drive automation, monitoring, and incident management on Kubernetes (Amazon EKS) with observability using Prometheus, Grafana, Dynatrace, and OpenSearch.
You will own SLI/SLO/SLA definitions, automate provisioning with Terraform, and lead on-call incident response, RCA, and cost optimization.
We are seeking a highly skilled Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems and applications. This role drives automation, monitoring, and incident management practices while operating Kubernetes platforms (Amazon EKS) and leveraging observability tools such as Prometheus, Grafana, Dynatrace, and OpenSearch to maintain high availability and operational excellence.