Site Reliability Engineer - ClickHouse

JobsinAustria

St. Pölten

Vor Ort

EUR 70.000 - 110.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

PostHog is seeking SREs to take deep ownership of production systems, especially a petabyte-scale ClickHouse deployment on AWS. The role centers on turning a fast-growing, stateful platform into a reliable, automated system, with provisioning, scaling, recovery, and self-healing automation at its core.

You'll work across databases and infra, reducing operational stress and building patterns that scale without proportionally increasing human effort.

Qualifikationen

  • Prior experience with ClickHouse or other OLAP databases.

Aufgaben

  • Managing large fleets of EC2‑based VMs, disks, and networking for data‑intensive workloads.
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response.
  • Working closely with ClickHouse engineers to turn database-level needs into infra-level solutions.
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation.
  • Participating in on‑call and incident response, with a strong focus on reducing incident frequency over time.
  • Designing and automating, not just responding to alerts.

Kenntnisse

ClickHouse
AWS
VMs
Terraform
Ansible
Linux
On-call

Jobbeschreibung

Who We're Looking For

We’re looking for people (EU/UK based) that like deep ownership of production systems, people that are not afraid of working with stateful infrastructure and love working in AWS, VMs, automation, and making messy systems reliable.

In General We Seek SREs Who Are
  • Enthusiastic drivers. We need proactive people that can fully own projects and get them done, and know to get help when needed. "Are we there yet?" is the wrong question.
  • Optimistic problem solvers. Things get hard here sometimes, whether it's scaling, shipping complex products, handling a stream of support requests, or trying to ship something that touches multiple teams. We need people who won't get disheartened, and will collaborate, iterate, and ship their way out of anything.
  • Grown ups. We’re an international bunch of weirdos, but one thing unites us: everyone is kind, considerate, and professional towards each other. This isn't about age or experience, it's about being low-ego, flexible, and respectful.
  • Genuine builders. PostHog is full of people who just love building stuff, people who would still be building software even if there wasn't a paycheck at the end. If this sounds like you, we should talk.
What You'll Be Doing

We run one of the largest self-managed ClickHouse installations on AWS, at petabyte scale, and we’re actively preparing it for the next 10–50× of growth. This role sits at the centre of that effort. You won’t be in a typical "keep the lights on" SRE role. The work is about turning a fast-growing, stateful system into a predictable, well-automated platform (provisioning, scaling, rebalancing, recovery). That means reducing operational stress, designing safe automation for data‑heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort. You’ll work on the kind of problems that only show up at large scale (petabytes of data, thousands of cores, constant ingestion).

  • Managing large fleets of EC2‑based VMs, disks, and networking for data‑intensive workloads.
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response.
  • Working closely with ClickHouse engineers to turn database‑level needs into infra‑level solutions.
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self‑healing automation.
  • Participating in on‑call and incident response, with a strong focus on making incidents rarer over time.
  • Designing and automating, not just responding to alerts.

You should join this team if you like deep ownership of production systems, and are not afraid of working with stateful infrastructure.

Requirements
  • Prior experience with ClickHouse or other OLAP databases.
  • Strong experience operating production infrastructure on AWS.
  • Hands‑on experience with VM‑based systems (EC2), not just managed PaaS.
  • Experience automating infrastructure using tools like Terraform, Ansible, or similar.
  • Solid understanding of Linux systems (disk, memory, networking, failure modes).
  • Experience supporting stateful systems (databases, queues, storage systems, etc.).
  • Ability to debug and reason about performance and reliability issues in production.
  • You’re comfortable owning systems end‑to‑end, including on‑call responsibilities.

You don’t need to be a ClickHouse expert on day one. We’ll teach you the database internals, but you do need to enjoy owning complex infrastructure.

We are committed to ensuring a fair and accessible interview process. If you need any accommodations or adjustments, please let us know.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer — AWS + ClickHouse, Stateful Infra
Site Reliability Engineer — AWS + ClickHouse, Stateful Infra

JobsinAustria • St. Pölten

Vor Ort
EUR 70.000 - 110.000
Senior Backend Engineer, Core APIs (SRE Focus)
Senior Backend Engineer, Core APIs (SRE Focus)

Jobgether • Lavamünd

Vor Ort
CHF 120.000 - 190.000
Fully remote working environment
Global, asynchronous collaboration
Exposure to modern tech stack: Go,Node
+2
Data Engineer
Data Engineer

Helu • Wien

Hybrid
EUR 70.000 - 95.000
Remote options
Quarterly Vienna weeks
Senior Backend Engineer
Senior Backend Engineer

Antithesis • Wien

Vor Ort
EUR 112.000 - 173.000
Senior Systems Engineer, Database
Senior Systems Engineer, Database

Antithesis • Wien

Vor Ort
EUR 55.000 - 75.000
Software Engineer (DevOps) - Site Reliability
Software Engineer (DevOps) - Site Reliability

Revolut • Österreich

Vor Ort
EUR 60.000 - 80.000
Senior Backend Engineer — Design High-Performance APIs
Senior Backend Engineer — Design High-Performance APIs

Antithesis • Wien

Vor Ort
EUR 112.000 - 173.000
Platform Engineer - Deployment Team
Platform Engineer - Deployment Team

Blackshark.ai GmbH • Graz

Vor Ort
EUR 60.000 - 70.000
Flexible working arrangements
Learning opportunities
Mental well-being programs
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Eversports • Wien

Hybrid
EUR 63.000 - 77.000
35h work week
Hybrid work (Office & Home)
MacBook provided
+2
Lead Platform Consultant - Openshift
Lead Platform Consultant - Openshift

SQUER • Wien

Hybrid
EUR 75.000 - 90.000
Flat hierarchies
Annual training budget
Hybrid work options
+4