Senior Cloud SRE — AI-Driven Uptime & Kubernetes

Clearwater Analytics, LLC

Boise (ID)

On-site

USD 130,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Clearwater Analytics is seeking a Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems. You will drive automation, monitoring, and incident management across Kubernetes platforms (Amazon EKS) and leverage observability tools to maintain high availability.

You will own SLI/SLO/SLA definitions, automate with Terraform, and lead on-call RCA. The role emphasizes AI-assisted investigations and collaboration with development teams for

Qualifications

  • 7+ years in Site Reliability Engineering or Platform Engineering.
  • Strong ownership of incident management and RCA for production systems.
  • Proven experience with Terraform and IaC in large-scale environments.
  • Hands-on AWS and EKS administration and operations.

Responsibilities

  • Design, build, and maintain highly available production systems.
  • Define and manage SLIs, SLOs, and SLAs to drive reliability.
  • Automate infrastructure provisioning and operations with Terraform.
  • Operate cloud-native platforms, including Amazon EKS.
  • Implement monitoring, logging, and alerting with Prometheus, Grafana, Dynatrace, OpenSearch.
  • Lead on-call rotation and RCA, driving continuous improvement.
  • Drive AI-assisted investigations and build AI-driven triage tooling.
  • Collaborate with development teams to improve reliability and deployment.
  • Build and maintain CI/CD pipelines (GitLab CI, Jenkins, GitHub Actions).
  • Perform capacity planning and cost optimization; ensure security and best practices.

Skills

Terraform
AWS
Kubernetes
Observability
Python/Go
CI/CD
Incident management

Education

Bachelor's degree in CS or related field

Tools

Prometheus
Grafana
Dynatrace
OpenSearch

Job description

Clearwater Analytics is seeking a Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems. You will drive automation, monitoring, and incident management across Kubernetes platforms (Amazon EKS) and leverage observability tools to maintain high availability.

You will own SLI/SLO/SLA definitions, automate with Terraform, and lead on-call RCA. The role emphasizes AI-assisted investigations and collaboration with development teams for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability & AI-Driven Incident Triage
Senior SRE: Cloud Reliability & AI-Driven Incident Triage

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000
Senior SRE: AI Cloud Infra, Kubernetes & Terraform (Remote)
Senior SRE: AI Cloud Infra, Kubernetes & Terraform (Remote)

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 140,000 - 190,000
Remote equipment stipend
Annual learning and development budget
Equity / Stock Options
+1
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)

Motion Recruitment • Chicago (IL)

On-site
USD 140,000 - 190,000
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

NDEAVOUR CONSULTING • United States

Hybrid
USD 120,000 - 150,000
Remote Office
Parking Space
Fun Office Space
+7
Senior Cloud SRE — AWS, Automation & Security
Senior Cloud SRE — AWS, Automation & Security

United States Digital Space LLC • Bellevue (CA)

On-site
USD 180,000 - 260,000
Amazing Benefits
Making Social Impact
Diversity, Equity & Inclusion
Senior Site Reliability Engineer – Cloud, Kubernetes & Automation
Senior Site Reliability Engineer – Cloud, Kubernetes & Automation

Socure • United States

On-site
USD 160,000 - 180,000
Senior SRE - AI-Powered Cloud Reliability (Remote)
Senior SRE - AI-Powered Cloud Reliability (Remote)

ServiceTitan, Inc. • Northern (KY)

Hybrid
USD 148,000 - 221,000
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7