Senior Site Reliability Engineer (m/f/d)

FACT-Finder Holding GmbH

Pforzheim

Vor Ort

EUR 80.000 - 110.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

FACT-Finder is transforming its product discovery platform with Kubernetes, Harvester, and GitOps. As a Senior Site Reliability Engineer, you will own SLOs/SLIs and steer incident response for reliability and performance across two products.

You will automate with Argo CD/Flux, help roll out auto-scaling, and contribute to observability, runbooks, and cost planning—spanning on-prem and cloud, with a hybrid work model.

Qualifikationen

  • Experience as an SRE, infrastructure, or production engineer in a SaaS or platform environment, or strong software/ops background with growth into SRE.
  • Solid understanding of SLOs, error budgets, incident management, and observability.
  • Hands-on experience with Kubernetes and cluster lifecycle upgrades.
  • Familiar with GitOps (Argo CD / Flux) and automation to reduce toil.
  • Experience with or interest in Harvester or comparable HCI/virtualization platforms.
  • Knowledge of auto-scaling primitives (HPA, VPA, KEDA) and capacity planning.
  • Fluent English; German is a plus.

Aufgaben

  • Define and own SLOs, SLIs, and error budgets across products for reliability and performance.
  • Drive incident response with blameless postmortems and follow-through.
  • Reduce manual work via automation and GitOps and build self-healing capabilities.
  • Support NG Search Operator development and auto-scaling rollouts.
  • Evolve observability with metrics, logs, traces, alerts, and runbooks.
  • Plan capacity and cost across on-prem and cloud, including burst scenarios.

Kenntnisse

Kubernetes
GitOps
Automation
Observability
SRE practices
Argo CD
Flux
Cluster lifecycle
Cost optimization
AI-assisted tooling

Tools

Kubernetes
Harvester
KEDA
OpenStack
vSphere/ESXi

Jobbeschreibung

IntroductionFACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. Both products are moving toward a modern, hybrid platform based on Kubernetes and Harvester - with the option to scale fully into the cloud in the mid-term. As a Senior Site Reliability Engineer (SRE), you make sure our systems stay fast, available, and scalable throughout this transformation. You work closely with the Hosting team and experienced engineers, and actively shape our journey toward a modern SaaS company.

Your mission
  • You define and own SLOs, SLIs, and error budgets across both products and make data-driven decisions on reliability and performance.
  • You drive incident response: fast detection, clear communication, blameless postmortems, and meaningful follow-through.
  • You consistently reduce manual work through automation and GitOps (e.g. Argo CD / Flux) and build out self-healing and self-service capabilities.
  • You support the development of an NG Search Operator (custom Kubernetes operator / CRDs) and the rollout of auto-scaling (HPA, VPA, KEDA, cluster autoscaler).
  • You evolve our observability - metrics, logs, traces, alerting, and runbooks that actually help on call.
  • You plan capacity and cost across on-premise (Frankfurt, Stockholm) and cloud - including burst scenarios into the public cloud.
  • You leverage AI tools to noticeably accelerate diagnosis, alerting, and operational workflows.
Your profile
  • Experience as an SRE, infrastructure, or production engineer in a SaaS or platform environment - or a strong software/operations background with a clear drive to grow into an SRE role.
  • Solid understanding of SLOs, error budgets, incident management, and observability.
  • Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns.
  • Experience with or strong interest in Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
  • Familiarity with GitOps (Argo CD / Flux), container storage (Longhorn, Ceph), and Kubernetes networking (load balancing, ingress).
  • Knowledge of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and capacity planning on-prem and in the cloud.
  • Understanding of networking in production-grade datacenters (incl. VLAN).
  • A strong automation instinct and a mindset to structurally eliminate toil.
  • Practical experience using AI tools in day-to-day operations.
  • Fluent English; German is a plus.
THE JOY OF WORKING WITH US
  • Impact from day one : Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud - with room to build things right.
  • AI-first mindset : We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.
Job Location

Berlin, Munich or Pforzheim (Hybrid)

About us

We are one of the leading Product Discovery Platforms for European eCommerce . Our search, navigation and recommendations handle billions of shopper queries every year - driving a combined GMV of more than €160 billion across our customers. These include Intersport, Spar, Douglas and more than 2,000 further B2B and B2C eCommerce companies across Europe. For over 20 years we have been an essential part of the complex eCommerce landscape, and with our AI-powered Product Discovery Experience (PDX) we are now entering the next phase - as a Private-Equity-owned company (Genui) , under a new McKinsey-shaped management team, with a clear growth and EBITDA plan. Offices in Pforzheim, Berlin, Munich and Stockholm.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

GoHiring GmbH • Deutschland

Hybrid
EUR 90.000 - 120.000
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

GoHiring GmbH • Deutschland

Hybrid
EUR 110.000 - 150.000
Competitive salary
Modern equipment
Learning budget
+1
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 140.000
Hybrid work model
Impact on product reliability
Competitive compensation
+1
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FactFinder • Berlin

Vor Ort
EUR 90.000 - 130.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

FACT-Finder • Pforzheim

Hybrid
EUR 110.000 - 170.000
Competitive salary
Modern equipment
Learning budget
+2
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Fact Finder • Berlin

Vor Ort
EUR 110.000 - 150.000
Wettbewerbsfähiges Gehalt
Weiterbildungsbudget
Regelmäßige Team-Events
+1
Customer Success Manager (all genders)
Customer Success Manager (all genders)

FactFinder • Berlin

Vor Ort
EUR 70.000 - 110.000
Großzügige Provision (OTE)
Team mit Training und Enablement
Arbeiten mit starken Marken & C‑Level‑
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

United States Digital Space LLC • Berlin

Hybrid
EUR 110.000 - 160.000
Wettbewerbsfähiges Gehalt
Moderne Ausstattung
Weiterbildungsbudget
+1
Staff Engineer - Manufacturing Systems & Reliability (f/m/d)
Staff Engineer - Manufacturing Systems & Reliability (f/m/d)

United States Digital Space LLC • Richen

Hybrid
EUR 90.000 - 130.000
Stock options
Ownership and impact
Hybrid work flexibility
+1