Team Lead - Site Reliability Engineering (all genders)

GoHiring GmbH

Deutschland

Vor Ort

EUR 110.000 - 150.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Competitive salary
Modern equipment
Learning budget
Team events

Zusammenfassung

FACT-Finder seeks a Team Lead Site Reliability Engineering to own reliability, scalability, and hosting costs as it modernizes toward Kubernetes on Harvester. You will lead the Hosting team, drive architecture decisions, and implement GitOps, observability, and incident management across on-prem and cloud.

You will shape a production-grade k8s platform, manage capacity and costs, and mentor engineers while exploring AI-driven operational insights. Fluency in English is required; German is a plus.

Qualifikationen

  • Strong background in infrastructure or platform engineering across on-premise and cloud.
  • Hands-on depth with Kubernetes in production: cluster lifecycle, upgrades, networking, storage, RBAC, observability, GitOps delivery.
  • Proven people leadership experience, excellent communication and stakeholder management skills.
  • Ideally practical experience with Harvester or comparable HCI/VM platforms.
  • Experience leading migration from bare metal/VMs to a k8s-based platform.
  • Comfort designing Kubernetes operators (custom controllers / CRDs).
  • Solid grasp of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and how they interact with capacity planning on-prem and in the cloud.

Aufgaben

  • Own the operational health of our hosting across on-premise and cloud – availability, performance, and incident management.
  • Drive modernization toward Kubernetes on Harvester: cluster topology, storage (Longhorn), networking, backup, and disaster recovery.
  • Build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.
  • Shape the NG Search Operator and solve auto-scaling for the current architecture.
  • Define our on-prem hybrid model: workloads, cloud bursting, latency, and cost control while keeping portability.
  • Own capacity planning and hosting cost and turn cost into a lever.
  • Lead and develop the Hosting team, set standards and ownership culture.
  • Make AI a core part of operations: diagnosis, automation, monitoring, and insight.

Kenntnisse

Kubernetes production
People leadership
GitOps tooling
RBAC & secrets
English fluent
Harvester familiarity

Tools

KubeVirt
vSphere/ESXi
OpenStack
Longhorn

Jobbeschreibung

Team Lead - Site Reliability Engineering (all genders)

FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. We are actively modernizing our hosting toward Kubernetes on Harvester – as an on-prem hybrid with the option to scale fully into the cloud in the mid-term. As Team Lead Site Reliability Engineering (all genders), you own the reliability, scalability, and cost of our hosting environments, drive this transformation end-to-end, and lead the team that delivers it.

Your mission
  • You own the operational health of our hosting across on-premise (Frankfurt, Stockholm) and cloud – availability, performance, and incident management.
  • You actively drive the modernization toward Kubernetes on Harvester: cluster topology, storage (Longhorn), networking (VLAN, load balancing, ingress), backup, and disaster recovery.
  • You build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.
  • You shape the NG Search Operator (custom Kubernetes operator) and solve auto-scaling (HPA, VPA, KEDA, cluster autoscaler) for the current architecture.
  • You concretely define our on-prem hybrid model: which workloads run where, how we burst into the cloud, how we keep latency and cost under control – while keeping the architecture portable enough for a future cloud-only move.
  • You own capacity planning and hosting cost and turn cost into a deliberate, managed lever.
  • You lead and develop our currently 4-person Hosting team, own performance and technical direction, and set the standards and ownership culture.
  • You make AI a core part of our operations: diagnosis, automation, monitoring, and insight.
Your profile
  • Strong background in infrastructure or platform engineering across on-premise and cloud.
  • Hands-on depth with Kubernetes in production: cluster lifecycle, upgrades, networking, storage, RBAC, observability, GitOps delivery.
  • Proven people leadership experience, excellent communication and stakeholder management skills.
  • Ideally practical experience with Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
  • Experience leading a real migration from bare metal / classic VMs to a k8s-based platform – including stateful workloads, storage migration, cutover, and rollback.
  • Comfort designing or operating Kubernetes operators (custom controllers / CRDs), ideally for stateful systems like search, databases, or streaming.
  • Solid grasp of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and how they interact with capacity planning on-prem and in the cloud.
  • Experience with on-prem hybrid architectures and owning reliability, capacity, and cost for production systems.
  • Hands-on fluency with AI tools in day-to-day operations.
  • Fluent English; German is a plus.
THE JOY OF WORKING WITH US
  • Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Leadership with real scope: You lead an established team and shape our platform in a decisive phase of our transformation.
  • Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.
  • Attractive benefits: Competitive salary, modern equipment, learning budget, and regular team events.
Job Location
About us

We are one of the leading Product Discovery Platforms for European eCommerce. Our search, navigation and recommendations handle billions of shopper queries every year — driving a combined GMV of more than €160 billion across our customers. These include Intersport, Spar, Douglas and more than 2,000 further B2B and B2C eCommerce companies across Europe.
For over 20 years we have been an essential part of the complex eCommerce landscape, and with our AI-powered Product Discovery Experience (PDX) we are now entering the next phase — as a Private-Equity-owned company (Genui), under a new McKinsey-shaped management team, with a clear growth and EBITDA plan. Offices in Pforzheim, Berlin, Munich and Stockholm.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

GoHiring GmbH • Deutschland

Hybrid
EUR 90.000 - 120.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

FACT-Finder • Pforzheim

Hybrid
EUR 110.000 - 170.000
Competitive salary
Modern equipment
Learning budget
+2
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Meyandy LLC • Berlin

Vor Ort
EUR 110.000 - 150.000
Competitive salary
Learning budget
Hybrid work model
+1
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Jackalope Digital LLC • Berlin

Hybrid
EUR 110.000 - 160.000
Competitive salary
Modern equipment
Learning budget
+2
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 140.000
Hybrid work model
Impact on product reliability
Competitive compensation
+1
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Hybrid work model
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Fact Finder • Berlin

Vor Ort
EUR 110.000 - 150.000
Wettbewerbsfähiges Gehalt
Weiterbildungsbudget
Regelmäßige Team-Events
+1
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

FactFinder • Berlin

Hybrid
Confidential
Wettbewerbsfähiges Gehalt
Moderne Ausstattung
Weiterbildungsbudget
+2
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

United States Digital Space LLC • Berlin

Hybrid
EUR 110.000 - 160.000
Wettbewerbsfähiges Gehalt
Moderne Ausstattung
Weiterbildungsbudget
+1