Senior Site Reliability Engineer: Cloud, Reliability & Automation

ClickHouse, Inc.

United Kingdom

Remote

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible work environment
Healthcare
Equity in the company
Time off
A USD$500 Home office setup
Global Gatherings

Job summary

ClickHouse, Inc. is expanding its central Site Reliability Engineering team. You will build and lead processes to ensure reliability, availability, and performance of our cloud infrastructure.

You will guide multiple engineering teams to design scalable, secure, and highly available distributed systems and own incident management, post-mortems, and continuous improvement of Cloud services. You will leverage Go and Python expertise to develop platforms and tools that optimize operational

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 8+ years of experience in Site Reliability Engineering or a related field.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Excellent understanding of distributed databases and SQL, particularly ClickHouse is a major plus.
  • Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • You are a strong problem solver with solid production debugging skills.
  • You are passionate about efficiency, availability, scalability, and data governance.
  • You thrive in a fast paced environment with ownership and accountability.
  • Excellent communication and interpersonal skills.

Responsibilities

  • Collaborate with engineering teams to design scalable, secure, highly available systems.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • Ensure monitoring and alerting across Data/Control Plane and Core components.
  • Lead incident response, post-mortems, and blameless analysis.
  • Improve reliability and performance of ClickHouse services.
  • Plan and drive Chaos initiatives across Engineering teams.
  • Manage on-call processes and incident escalation practices.

Skills

Go
Python
AWS
Azure
GCP
ClickHouse
SQL

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

Kubernetes
Docker Swarm
Ansible
Terraform
Puppet

Job description

ClickHouse, Inc. is expanding its central Site Reliability Engineering team. You will build and lead processes to ensure reliability, availability, and performance of our cloud infrastructure.

You will guide multiple engineering teams to design scalable, secure, and highly available distributed systems and own incident management, post-mortems, and continuous improvement of Cloud services. You will leverage Go and Python expertise to develop platforms and tools that optimize operational

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer- Remote
Senior Site Reliability Engineer- Remote

ClickHouse, Inc. • United Kingdom

Remote
GBP 90,000 - 130,000
Flexible work environment
Healthcare
Equity in the company
+3
Lead Site Reliability Engineer | Kubernetes & Automation
Lead Site Reliability Engineer | Kubernetes & Automation

Factset • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage
Free lunch in the office (Mon–Fri)
Employee social events and sports
Senior Site Reliability Engineer - Kubernetes & Automation
Senior Site Reliability Engineer - Kubernetes & Automation

20035 FactSet Europe Limited • Greater London

Hybrid
GBP 90,000 - 150,000
Senior Platform SRE: Multi-Cloud Reliability & Resilience
Senior Platform SRE: Multi-Cloud Reliability & Resilience

Elastic • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage for you and family
Flexible location & schedule
Generous vacation days
+3
Senior Cloud SRE: Multi-Cloud Reliability & Automation
Senior Cloud SRE: Multi-Cloud Reliability & Automation

Elasticsearch B.V. • United Kingdom

Remote
GBP 75,000 - 140,000
Health coverage
Flexible locations and schedules
Generous vacation days
+4
Site Reliability Engineer - Live Ops & Cloud Resilience
Site Reliability Engineer - Live Ops & Cloud Resilience

World Wrestling Entertainment, Inc. • Greater London

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer – Cloud, Observability & Resilience
Site Reliability Engineer – Cloud, Observability & Resilience

IMG • Greater London

Hybrid
GBP 70,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Senior Site Reliability Engineer - Cloud Observability & Automation
Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior SRE: Cloud-Native Reliability Lead (Remote)
Senior SRE: Cloud-Native Reliability Lead (Remote)

Spotme • United Kingdom

Remote
GBP 90,000 - 130,000
Work from anywhere
MacBook setup budget
Inclusive culture
+5