Senior Site Reliability Engineer / SRE - Kubernetes & Hybrid Cloud (m/f/d)

FactFinder

Berlin

Hybrid

EUR 90.000 - 140.000

Vollzeit

Vor 5 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Hybrid work model (3 office days/week)

Zusammenfassung

FactFinder in Berlin is seeking an experienced SRE to join our growing team. You'll help build a modern private cloud platform on Kubernetes and Harvester, with on‑prem by default and elastic bursts into the public cloud, while collaborating with developers and two system admins in Pforzheim.

You'll define SLOs/SLIs, lead incident response, automate toil with GitOps, and contribute to a stateful search Kubernetes operator.

Qualifikationen

  • Kubernetes in production—set up and maintained clusters on own servers.
  • Lived SRE practice: SLOs, error budgets, incident management, on-call.
  • Hands-on GitOps/infrastructure automation; Argo CD or Flux a strong plus.
  • Strong observability skills—metrics, logs, traces, alerting.
  • Automation mindset; you prefer fixing root causes over applying quick workarounds.
  • Collaborative, enabling mindset—serving developers with fast, well-communicated reliability decisions.

Aufgaben

  • Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions.
  • Lead incident response end-to-end: fast detection, clear communication, blameless postmortems.
  • Eliminate toil through automation and GitOps; evolve observability across two stacks.
  • Help build our Kubernetes operator (CRDs) for stateful search clusters, with self-healing and safe upgrades.
  • Plan capacity, performance and cost across on-prem and cloud, aligning with merchant loads and AI-assisted diagnosis.

Kenntnisse

Kubernetes production
SRE practice
GitOps
Argo CD/Flux
Observability
Automation
Team collaboration
Incident management/on-call

Tools

Kubernetes
Harvester
KubeVirt
Prometheus
Grafana
Longhorn
Ceph
kubeadm
RKE2
k3s

Jobbeschreibung

Introduction

At a glance

Location & workmodel: Berlin, hybrid

Tech stack: Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph

Team: A growing SRE team - you report to our CTPO for now and to the Team Lead SRE we’re hiring next; two system administrators in Pforzheim run the physical hardware

Process: Intro call - take-home task (~2h) - 90-min tech interview with our developers - leadership conversation - meet the team

Languages: Fluent English required; German is a plus, nota must

Why this role is special

Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester - on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams - and you won't do it alone, a Team Lead SRE hire is coming next.

SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct - our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.

Your first 90 days

You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic - SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.

Your mission
  • Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions
  • Lead incident response end-to-end: fast detection, clear communication, blameless postmortems - and reduce whole classes of incidents structurally, not case by case
  • Eliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks
  • Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable - and roll out the auto-scaling (HPA/VPA, KEDA, clusterautoscaler) today's architecture makes hard
  • Plan capacity, performance and cost across on-premises and cloud - including the large catalogue and peak-season loads our merchants care about - and use AI tools wherever they measurably speed up diagnosis and operations
Your profile
Must-haves
  • Kubernetes in production - built, not just used: you've set up and maintained clusters on your own servers (e.g. kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades - managed-only experience isn't enough for this role
  • Lived SRE practice: SLOs, error budgets, incident management, on-call
  • Hands-on experience with GitOps or comparable infrastructure/deployment automation - experience with Argo CD or Flux is a strong plus
  • Solid observability skills - metrics, logs, traces, alerting that people trust
  • A strong automation instinct - you'd rather fix a problem's cause than repeat its workaround
  • A collaborative, enabling mindset - you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solution
Nice-to-haves (genuinely optional - we'll teach you the rest)
  • Harvester, KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms
  • Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
  • Auto-scaling (HPA, VPA, KEDA, clusterautoscaler) and capacity/cost planning
  • Experience building Kubernetes operators/CRDs
  • German language skills
  • Certifications (CKA, CKS) are welcome but no substitute for hands-on experience - in the tech interview we'll ask about what you've actually built and operated.
THE JOY OF WORKING WITH US
  • Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud - with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model three office days per week with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment wher
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FactFinder • Berlin

Hybrid
Confidential
Hybrid work model
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 125.000
Hybrid work model
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

Meyandy LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Hybrid work model
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FactFinder • Berlin

Hybrid
Confidential
Hybrides Arbeitsmodell
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FactFinder • Berlin

Vor Ort
EUR 90.000 - 130.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

FACT-Finder • Pforzheim

Vor Ort
EUR 110.000 - 170.000
Competitive salary
Modern equipment
Learning budget
+2
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

Jackalope Digital LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Hybrid work model
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Meyandy LLC • Berlin

Vor Ort
EUR 110.000 - 150.000
Competitive salary
Learning budget
Hybrid work model
+1
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Fact Finder • Berlin

Vor Ort
EUR 110.000 - 150.000
Wettbewerbsfähiges Gehalt
Weiterbildungsbudget
Regelmäßige Team-Events
+1
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Jackalope Digital LLC • Berlin

Hybrid
EUR 110.000 - 160.000
Competitive salary
Modern equipment
Learning budget
+2