Sr. Site Reliability Engineer

Clearwater Analytics, LLC

Boise (ID)

On-site

USD 130,000 - 170,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Clearwater Analytics is seeking a Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems. You will drive automation, monitoring, and incident management across Kubernetes platforms (Amazon EKS) and leverage observability tools to maintain high availability.

You will own SLI/SLO/SLA definitions, automate with Terraform, and lead on-call RCA. The role emphasizes AI-assisted investigations and collaboration with development teams for

Qualifications

  • 7+ years in Site Reliability Engineering or Platform Engineering.
  • Strong ownership of incident management and RCA for production systems.
  • Proven experience with Terraform and IaC in large-scale environments.
  • Hands-on AWS and EKS administration and operations.

Responsibilities

  • Design, build, and maintain highly available production systems.
  • Define and manage SLIs, SLOs, and SLAs to drive reliability.
  • Automate infrastructure provisioning and operations with Terraform.
  • Operate cloud-native platforms, including Amazon EKS.
  • Implement monitoring, logging, and alerting with Prometheus, Grafana, Dynatrace, OpenSearch.
  • Lead on-call rotation and RCA, driving continuous improvement.
  • Drive AI-assisted investigations and build AI-driven triage tooling.
  • Collaborate with development teams to improve reliability and deployment.
  • Build and maintain CI/CD pipelines (GitLab CI, Jenkins, GitHub Actions).
  • Perform capacity planning and cost optimization; ensure security and best practices.

Skills

Terraform
AWS
Kubernetes
Observability
Python/Go
CI/CD
Incident management

Education

Bachelor's degree in CS or related field

Tools

Prometheus
Grafana
Dynatrace
OpenSearch

Job description

Job Summary

We are seeking a highly skilled Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems and applications. This role drives automation, monitoring, and incident management practices while operating Kubernetes platforms (Amazon EKS) and leveraging observability tools such as Prometheus, Grafana, Dynatrace, and OpenSearch to maintain high availability and operational excellence.

Key Responsibilities
  • Design, build, and maintain highly available, scalable, and reliable production systems.
  • Define and manage SLIs, SLOs, and SLAs to drive system reliability.
  • Automate infrastructure provisioning and operations using Terraform (IaC).
  • Operate and manage cloud-native platforms, including Amazon EKS.
  • Implement and maintain monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, and OpenSearch.
  • Lead incident management - on-call rotation, production troubleshooting, and root cause analysis (RCA).
  • Drive AI-assisted investigations as a core part of incident response, and build and maintain the prompts, integrations, and guardrails that make AI-driven triage and RCA reliable.
  • Improve system reliability through automation, self-healing mechanisms, and performance tuning.
  • Collaborate with development teams to improve application reliability, scalability, and deployment processes.
  • Build and maintain CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions) for fast, reliable software delivery.
  • Perform capacity planning and cost optimization for infrastructure and services.
  • Ensure security, compliance, and best practices across infrastructure and applications.
Required Qualifications
  • Bachelor's degree in computer science or a related field, or equivalent practical experience.
  • 7+ years in Site Reliability Engineering or Platform Engineering.
  • Proven ownership of incident management, on-call support, and root cause analysis (RCA) for production systems.
  • Strong expertise in Terraform and Infrastructure as Code.
  • Hands-on experience with AWS and EKS.
  • Strong understanding of monitoring, logging, and observability (Prometheus, Grafana, Dynatrace, OpenSearch).
  • Proficiency in Python, Java, Go, or Bash.
  • Experience with Agile development and CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions).
  • Strong problem-solving, documentation, and communication skills.
  • Proven ability to troubleshoot effectively in high-pressure production environments.
  • Experience with autoscaling, performance tuning, and cost optimization.
  • Familiarity with AI-assisted automation tools and a track record of using them to reduce toil and improve reliability.
Preferred Skills
  • Docker and Linux administration.
  • Build systems and dependency management (Maven, Gradle, npm).
  • Additional AWS services: Cognito, WAF, Elasticsearch, SNS, SQS, S3, Systems Manager.
  • Database infrastructure knowledge (RDS, MySQL, SQL Server).
  • Cloud or Kubernetes certifications.

Thank you for your interest in a career with Clearwater! Clearwater Analytics (NYSE: CWAN) is transforming investment management with the industry's most comprehensive cloud-native platform for institutional investors across global public and private markets. While legacy systems create risk, inefficiency, and data fragmentation, Clearwater's single-instance, multi-tenant architecture delivers real-time data and AI-driven insights throughout the investment lifecycle. The platform eliminates information silos by integrating portfolio management, trading, investment accounting, reconciliation, regulatory reporting, performance, compliance, and risk analytics in one unified system. Serving leading insurers, asset managers, hedge funds, banks, corporations, and governments, Clearwater supports over $8.8 trillion in assets globally. Learn more at www.clearwateranalytics.com. Studies have shown that women and people of color are less likely to apply to jobs unless they meet every single qualification. We are dedicated to building a diverse, inclusive and authentic workplace, so if you're excited about this role but your past experience doesn't align perfectly with the job description,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Clearwater Analytics • Chicago (IL)

On-site
USD 95,000 - 123,000
Health/vision/dental insurance
401(k)
PTO
+3
Sr. Subject Matter Expert - Implementations
Sr. Subject Matter Expert - Implementations

Clearwater Analytics, LLC • Boise (ID)

On-site
USD 110,000 - 135,000
Client Servicing Analyst
Client Servicing Analyst

Clearwater Analytics, LLC • Illinois

On-site
USD 54,000 - 70,000
health/vision/dental insurance
401(k)
PTO
+3
S&M / Ops Support Intern
S&M / Ops Support Intern

Clearwater Analytics • Boise (ID)

On-site
USD 42,000 - 64,000
Data Flow Engineer
Data Flow Engineer

Clearwater Analytics • Boise (ID), Northern (KY)

On-site
USD 70,000 - 100,000
Software Development Engineer II
Software Development Engineer II

Clearwater Analytics, LLC • Boise (ID)

On-site
USD 95,000 - 135,000
Data Flow Engineer I
Data Flow Engineer I

Clearwater Analytics • Boise (ID), Northern (KY)

On-site
USD 90,000 - 120,000
Subject Matter Expert - Implementations
Subject Matter Expert - Implementations

Clearwater Analytics, LLC • Boise (ID)

On-site
USD 110,000 - 140,000
Senior Financial Analyst
Senior Financial Analyst

Clearwater Analytics • Chicago (IL)

On-site
USD 85,000 - 120,000
Health/vision/dental insurance
401(k)
PTO
+3
Implementation Project Manager
Implementation Project Manager

Clearwater Analytics • New York (NY), Northern (KY)

On-site
USD 102,000 - 144,000
Health/vision/dental insurance
401(k)
PTO
+3