Remote Senior SRE: Cloud Reliability & Incident Response

ClickHouse, Inc.

Canada

Hybrid

CAD 140,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible work environment
Healthcare
Equity in the company
Time off
A USD$500 Home office setup
Global Gatherings

Job summary

ClickHouse, Inc. is expanding its central Site Reliability Engineering team in Canada. You will build and lead processes to ensure reliability, availability, scalability, and performance of our cloud infrastructure and guide cross-functional teams to design resilient systems.

You will own incident management, post-mortem analysis, and drive continuous improvements, leveraging Go or Python to develop platforms that boost operational efficiency and reliability at scale.

Qualifications

  • 8+ years of Site Reliability Engineering or a related field.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Experience with container orchestration: Kubernetes or Docker Swarm.
  • Experience with automation/configuration tools: Ansible, Terraform, or Puppet.

Responsibilities

  • Collaborate with engineering teams to design scalable, secure, and highly available systems.
  • Establish and manage SLOs/SLAs for ClickHouse Cloud.
  • Ensure monitoring and alerting across all infrastructure components.
  • Improve incident response processes and post-mortem analysis.
  • Drive reliability and performance improvements for ClickHouse services.
  • Lead Chaos initiatives and manage on-call processes.

Skills

Go
Python
SRE experience
Communication

Education

Bachelor's or Master's in CS or related field

Tools

Kubernetes
Docker Swarm
Ansible
Terraform
Puppet

Job description

ClickHouse, Inc. is expanding its central Site Reliability Engineering team in Canada. You will build and lead processes to ensure reliability, availability, scalability, and performance of our cloud infrastructure and guide cross-functional teams to design resilient systems.

You will own incident management, post-mortem analysis, and drive continuous improvements, leveraging Go or Python to develop platforms that boost operational efficiency and reliability at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff SRE: Cloud Reliability & Scale on GCP
Staff SRE: Cloud Reliability & Scale on GCP

SoundHound AI • Toronto

On-site
CAD 140,000 - 180,000
Equity
Healthcare
Paid time off
Senior Cloud Software Engineer - Efficiency Engineering
Senior Cloud Software Engineer - Efficiency Engineering

ClickHouse • Canada

Remote
CAD 140,000 - 190,000
Flexible work environment
Healthcare
Equity
+3
Senior SRE: Cloud, CI/CD & Incident Leadership (Hybrid)
Senior SRE: Cloud, CI/CD & Incident Leadership (Hybrid)

Morningstar Credit Ratings, LLC • Toronto

On-site
CAD 90,000 - 133,000
Hybrid work environment
Flexible benefits
Staff Site Reliability Engineer Toronto, Canada (remote) •
Staff Site Reliability Engineer Toronto, Canada (remote) •

SoundHound Inc. • Toronto

Hybrid
CAD 140,000 - 200,000
Equity
Healthcare
Paid time off
+1
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Staff Incident Command & Reliability Engineer
Staff Incident Command & Reliability Engineer

IBM • Markham

On-site
CAD 120,000 - 180,000
Senior SRE – AI-Driven Cloud & Reliability Leader
Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius • Vancouver

Hybrid
CAD 115,000 - 145,000
SRE Manager — Remote Lead, Cloud Platform & DX
SRE Manager — Remote Lead, Cloud Platform & DX

Lightspeed Commerce, Inc. • Montreal (administrative region)

On-site
CAD 150,000 - 190,000
Flexible PTO
Equity options
Pension contributions
+5
Senior SRE: Linux & Cloud Infrastructure
Senior SRE: Linux & Cloud Infrastructure

Atlantis IT Group • Montreal

On-site
CAD 80,000 - 100,000
Senior Cloud Reliability Engineer (AWS/Kubernetes)
Senior Cloud Reliability Engineer (AWS/Kubernetes)

Enverus • Calgary

On-site
CAD 120,000 - 180,000