(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)

STACKIT

Heilbronn

Vor Ort

EUR 90.000 - 120.000

Vollzeit

Vor 8 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Schwarz Digits is seeking a Site Reliability Engineer to strengthen our distributed systems in production. You will collaborate with development teams to shorten MTTD, design robust CI/CD pipelines, and build automation in Go while applying reliability patterns and shift-left practices.

You will lead incident responses during on-call rotations and help optimize performance of the Control Plane with a keen focus on scalability, observability, and post-mortem learning.

Qualifikationen

  • 3+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering.
  • Deep knowledge of Kubernetes Control Plane internals and etcd.
  • Production-grade Go coding for automation and tooling.

Aufgaben

  • Collaborate with development teams to shorten time-to-detect and ensure services meet SLOs.
  • Improve time-to-mitigation by crafting playbooks and dashboards for responders.
  • Act as a reliability consultant, teaching teams reliability patterns and shift-left practices.
  • Design and refine CI/CD pipelines for Canary and Blue/Green deployments.
  • Analyze scalability of the Control Plane, addressing bottlenecks in distributed systems.
  • Participate in on-call rotation, leading incident responses and blameless post-mortems.

Kenntnisse

Kubernetes
Go
IaC
Linux
PostgreSQL
Redis
Kafka
NATS

Jobbeschreibung

Schwarz Digits creates the technological foundation for digital sovereignty in Europe. As the IT and digital division of the Schwarz Group, we develop and manage the IT infrastructures for the retail divisions Lidl and Kaufland, as well as Schwarz Production and PreZero. At the same time, we operate as an independent provider in the external market to support companies across Europe in their digital transformation. We bundle our core services in the areas of Cloud, Cyber Security, Data & AI, Communication, and Workspace.

Join us and contribute to digital sovereignty in Europe. With us, you will work at the intersection of agility and security: You will benefit from fast decision-making processes, enjoy genuine creative freedom in your projects, and be able to build upon the stable foundation of the Schwarz Group.

Your tasks
  • You collaborate closely with development teams to shorten time-to-detect intervals by enhancing our monitoring and alerting infrastructure and ensuring our services adhere to defined SLOs.
  • Your work is critical in continuously optimizing our time-to-mitigation; you achieve this by creating clear playbooks, designing dashboards for first responders, and ensuring our telemetry data (logs and metrics) is comprehensive.
  • You act as a reliability consultant to development teams, educating them on reliability patterns and helping them "shift left" to foster a shared responsibility model.
  • You design and refine development practices, including CI/CD pipelines, to support progressive delivery strategies such as Canary releases and Blue/Green deployments.
  • You proactively analyze and optimize the scalability of the Control Plane, addressing bottlenecks in distributed consensus, database throughput, and kernel-level networking.
  • You participate in a compensated on-call rotation, leading incident responses and facilitating blameless post-mortems and Root Cause Analyses.
Your profile
  • You bring 3+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a specific focus on operating large-scale distributed systems in production.
  • You possess expert-level knowledge of Kubernetes Control Plane internals, including the API Server, Controller Manager, Scheduler, and etcd.
  • You demonstrate proficiency in Go and write production-grade code to build automation tools, Kubernetes Operators, or glue code that integrates disparate systems.
  • You hold deep experience with Infrastructure as Code and container infrastructure, alongside proficiency in Linux system internals (kernel tuning, memory management) and networking (TCP/IP, CNI, Load Balancers, eBPF).
  • You bring experience in operating datastores (e.g., PostgreSQL, Redis) and messaging systems (e.g., Kafka, NATS) in scalable environments.
  • You run towards fires to learn from them, you automate yourself out of a job, and you believe that hope is not a strategy.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)

Schwarz Digits • Heilbronn

Vor Ort
EUR 55.000 - 75.000
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)

Schwarz Dienstleistung KG • Heilbronn

Vor Ort
EUR 60.000 - 80.000
Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/f/d)
Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/f/d)

STACKIT • Bad Friedrichshall

Vor Ort
EUR 90.000 - 130.000
(Senior) Platform Engineer - STACKIT Control Plane Platform (m/w/d)
(Senior) Platform Engineer - STACKIT Control Plane Platform (m/w/d)

Schwarz Dienstleistung KG • Heilbronn

Vor Ort
EUR 75.000 - 110.000
Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/w/d)
Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/w/d)

Schwarz Digits • Bad Friedrichshall

Vor Ort
EUR 65.000 - 85.000
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

STACKIT • Heilbronn

Vor Ort
EUR 70.000 - 110.000
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

Schwarz Digits • Heilbronn

Vor Ort
EUR 55.000 - 75.000
(Junior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Junior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

STACKIT • Heilbronn

Vor Ort
EUR 70.000 - 100.000
Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/f/d)
Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/f/d)

Schwarz Digits • Bad Friedrichshall

Vor Ort
EUR 55.000 - 75.000
(Senior) Software Engineer - STACKIT Control Plane Platform (m/w/d)
(Senior) Software Engineer - STACKIT Control Plane Platform (m/w/d)

STACKIT • Heilbronn

Vor Ort
EUR 70.000 - 110.000