Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/f/d)

STACKIT

Bad Friedrichshall

Vor Ort

EUR 90.000 - 130.000

Vollzeit

Vor 7 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Schwarz Digits seeks a Senior Engineer to build and own a real-time cloud messaging backbone. You will architect and scale a high-throughput, multi-datacenter message platform used by dozens of product teams, ensuring reliability and low latency.

You will lead deployments of Kafka/Solace/NATS clusters, drive operator-based Kubernetes workflows, and implement robust monitoring while enabling developers with self-service tooling.

Qualifikationen

  • Experience in managing large-scale distributed systems in production.
  • Hands-on administrative experience with enterprise brokers like Apache Kafka or Solace.
  • Experience in managing infrastructure on Kubernetes using the Operator Pattern.
  • Managing virtual machines using Ansible.
  • Fluency in Python, Go or Bash; strong Linux performance and networking knowledge.
  • Understanding of event-driven architecture and streaming concepts.

Aufgaben

  • Design, deploy, and manage highly available, distributed messaging clusters across data centers.
  • Ensure reliability, performance, and fault tolerance with DR and failover strategies; tune OS for low latency.
  • Automate provisioning, scaling, and configuration of messaging clusters.
  • Build monitoring, alerting, and logging dashboards for health and latency.
  • Define best practices and build a self-service platform for internal integrations.

Kenntnisse

Distributed systems
Apache Kafka
Solace
NATS
Kubernetes
Ansible
Python
Go
Bash
Linux performance tuning
Networking
Event-driven architecture
English communication
German communication

Tools

Apache Kafka
Solace
NATS
Kubernetes
Ansible

Jobbeschreibung

Schwarz Digits creates the technological foundation for digital sovereignty in Europe. As the IT and digital division of the Schwarz Group, we develop and manage the IT infrastructures for the retail divisions Lidl and Kaufland, as well as Schwarz Production and PreZero. At the same time, we operate as an independent provider in the external market to support companies across Europe in their digital transformation. We bundle our core services in the areas of Cloud, Cyber Security, Data & AI, Communication, and Workspace.

Join us and contribute to digital sovereignty in Europe. With us, you will work at the intersection of agility and security: You will benefit from fast decision-making processes, enjoy genuine creative freedom in your projects, and be able to build upon the stable foundation of the Schwarz Group.

We are looking for a Senior Engineer to build, scale, and own the central nervous system of our cloud infrastructure: a highly resilient, high-throughput message and event platform. As our engineering organization scales rapidly, we are transitioning to a real-time, event-driven architecture to ensure seamless communication between the control plane components of all our products. You will empower dozens of product teams by providing an outstanding developer experience, enabling them to seamlessly publish and consume millions of events per day.

Your Tasks
  • You design, deploy, and manage highly available, distributed message broker clusters (such as Apache Kafka, Solace, or NATS) across multiple data centers.
  • You ensure the reliability, performance, and fault tolerance of the messaging infrastructure by implementing robust disaster recovery and failover strategies and tune operating system configurations for low-latency delivery.
  • You automate the provisioning, scaling, and configuration of messaging clusters.
  • You build comprehensive monitoring, alerting, and logging dashboards to track cluster health, throughput, and latency.
  • You define best practices for application developers and build a self-service platform that makes it easy for internal teams to independently configure their integrations.
Your Profile
  • You bring solid experience in managing large-scale distributed systems in production, coming from a background like Site Reliability Engineering or Platform Engineering.
  • You have deep, hands-on administrative experience with enterprise brokers like Apache Kafka or Solace, and you bring experience in managing infrastructure on Kubernetes using the Operator Pattern as well as managing virtual machines using tools like Ansible.
  • You bring fluency in coding with Python, Go, or Bash, alongside a strong understanding of Linux performance tuning and networking protocols (such as Transmission Control Protocol/Internet Protocol or Domain Name System).
  • Ideally, you have a profound understanding of event-driven architecture patterns and event streaming concepts to help design scalable, real-time data pipelines.
  • Your English, ideally combined with German, is the basis for successful communication in our international, agile teams.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/w/d)
Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/w/d)

Schwarz Digits • Bad Friedrichshall

Vor Ort
EUR 65.000 - 85.000
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)

STACKIT • Heilbronn

Vor Ort
EUR 90.000 - 120.000
Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/f/d)
Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/f/d)

Schwarz Digits • Bad Friedrichshall

Vor Ort
EUR 55.000 - 75.000
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)

Schwarz Digits • Heilbronn

Vor Ort
EUR 55.000 - 75.000
Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/w/d)
Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/w/d)

STACKIT • Bad Friedrichshall

Vor Ort
EUR 70.000 - 100.000
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

STACKIT • Heilbronn

Vor Ort
EUR 70.000 - 110.000
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

Schwarz Digits • Heilbronn

Vor Ort
EUR 55.000 - 75.000
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

Schwarz Dienstleistung KG • Heilbronn

Vor Ort
EUR 60.000 - 80.000
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/w/d)

Schwarz Dienstleistung KG • Heilbronn

Vor Ort
EUR 60.000 - 80.000
(Senior) Customer Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
(Senior) Customer Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)

STACKIT • Heilbronn

Vor Ort
EUR 90.000 - 130.000