Senior SRE, Kubernetes Control Plane & Cloud Reliability

Schwarz Dienstleistung KG

Germany (OH)

On-site

USD 104,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading cloud provider is seeking an experienced Site Reliability Engineer (SRE) to join their Products division. Your role involves guiding system architecture, optimizing databases, and improving service reliability. Collaborating closely with development teams, you'll enhance monitoring systems and support CI/CD practices. Ideal candidates have over 3 years of experience in SRE or DevOps, with strong expertise in Kubernetes and Go. This offer comes with opportunities to influence large-scale systems in an innovative environment.

Qualifications

  • 3+ years in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Expertise with Kubernetes Control Plane internals.
  • Proficiency in Go and production-grade code development.

Responsibilities

  • Collaborate with development teams to enhance monitoring and alerting.
  • Design dashboards and playbooks for incident management.
  • Participate in on-call rotation and incident responses.

Skills

Site Reliability Engineering
Kubernetes
Go
Infrastructure as Code
Linux internals
Networking

Job description

Select how often (in days) to receive an alert:

Schwarz Digits creates the technological foundation for digital sovereignty in Europe. As the IT and digital division of the Schwarz Group, we develop and manage the IT infrastructures for the retail divisions Lidl and Kaufland, as well as Schwarz Production and PreZero. At the same time, we operate as an independent provider in the external market to support companies across Europe in their digital transformation. We bundle our core services in the areas of Cloud, Cyber Security, Data & AI, Communication, and Workspace.

Join us and contribute to digital sovereignty in Europe. With us, you will work at the intersection of agility and security: You will benefit from fast decision-making processes, enjoy genuine creative freedom in your projects, and be able to build upon the stable foundation of the Schwarz Group.

  • You collaborate closely with development teams to shorten time-to-detect intervals by enhancing our monitoring and alerting infrastructure and ensuring our services adhere to defined SLOs.
  • Your work is critical in continuously optimizing our time-to-mitigation; you achieve this by creating clear playbooks, designing dashboards for first responders, and ensuring our telemetry data (logs and metrics) is comprehensive.
  • You act as a reliability consultant to development teams, educating them on reliability patterns and helping them "shift left" to foster a shared responsibility model.
  • You design and refine development practices, including CI/CD pipelines, to support progressive delivery strategies such as Canary releases and Blue/Green deployments.
  • You proactively analyze and optimize the scalability of the Control Plane, addressing bottlenecks in distributed consensus, database throughput, and kernel-level networking.
  • You participate in a compensated on-call rotation, leading incident responses and facilitating blameless post-mortems and Root Cause Analyses.
Your profile
  • You bring 3+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a specific focus on operating large-scale distributed systems in production.
  • You possess expert-level knowledge of Kubernetes Control Plane internals, including the API Server, Controller Manager, Scheduler, and etcd.
  • You demonstrate proficiency in Go and write production-grade code to build automation tools, Kubernetes Operators, or glue code that integrates disparate systems.
  • You hold deep experience with Infrastructure as Code and container infrastructure, alongside proficiency in Linux system internals (kernel tuning, memory management) and networking (TCP/IP, CNI, Load Balancers, eBPF).
  • You bring experience in operating datastores (e.g., PostgreSQL, Redis) and messaging systems (e.g., Kafka, NATS) in scalable environments.
  • You run towards fires to learn from them, you automate yourself out of a job, and you believe that hope is not a strategy.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

(Senior) Site Reliability Engineer - STACKIT Control Plane (m/f/d)
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/f/d)

Schwarz Dienstleistung KG • Germany (OH)

On-site
USD 104,000 - 150,000
Senior SRE: Kubernetes Reliability & GitOps Lead
Senior SRE: Kubernetes Reliability & GitOps Lead

SysEleven GmbH • Germany (OH)

On-site
USD 100,000 - 130,000
STACKIT Platform Services - Observability Suite - Engineering Manager (m/f/d)
STACKIT Platform Services - Observability Suite - Engineering Manager (m/f/d)

Schwarz Dienstleistung KG • Germany (OH)

On-site
USD 120,000 - 150,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Kubernetes Platform Engineer Lead
Senior Kubernetes Platform Engineer Lead

TechDigital Group • Georgia

Hybrid
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Lead Kubernetes SRE
Lead Kubernetes SRE

TechDigital Group • Minneapolis (MN)

On-site
USD 80,000 - 120,000