(Senior) Site Reliability Engineer - STACKIT Control Plane (m/f/d)

Schwarz Dienstleistung KG

Germany (OH)

On-site

USD 104,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading cloud provider is seeking an experienced Site Reliability Engineer (SRE) to join their Products division. Your role involves guiding system architecture, optimizing databases, and improving service reliability. Collaborating closely with development teams, you'll enhance monitoring systems and support CI/CD practices. Ideal candidates have over 3 years of experience in SRE or DevOps, with strong expertise in Kubernetes and Go. This offer comes with opportunities to influence large-scale systems in an innovative environment.

Qualifications

  • 3+ years in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Expertise with Kubernetes Control Plane internals.
  • Proficiency in Go and production-grade code development.
  • Strong experience with Infrastructure as Code and container infrastructure
  • Knowledge of Linux kernel tuning, memory management, and networking (TCP/IP, CNI, Load Balancers, eBPF)
  • Experience operating datastores (PostgreSQL, Redis) and messaging systems (Kafka, NATS)

Responsibilities

  • Collaborate with development teams to enhance monitoring and alerting.
  • Design dashboards and playbooks for incident management.
  • Participate in on-call rotation and incident responses.
  • Participate in compensated on-call rotation and conduct blameless post-mortems
  • Push reliability practices and educate teams on shift-left reliability

Skills

Site Reliability Engineering
Kubernetes
Go
Infrastructure as Code
Linux internals
Networking

Tools

PostgreSQL
Redis
Kafka
NATS
etcd
Linux

Job description

You want to take the cloud world by storm with us STACKITEERs and shape the future of Europe with us? That's great! Then STACKIT is the right place for you. Our vision is ambitious: An independent Europe - digital, leading. As a cloud and colocation provider, we are building the secure infrastructure for this. With our server locations exclusively in Germany and Austria, we offer both the Schwarz Group, to which we belong, and external customers a European alternative to international cloud providers and support our customers holistically with individual solutions.

As a dedicated STACKITEER, you are part of the STACKIT Products division. This is where our products and services are developed, tested and improved.

As an SRE for the STACKIT Control Plane, you guide system architecture by operating at the intersection of development and systems engineering. Together with the development team, you design, build, and run large-scale systems that are inherently scalable and reliable. Your challenges range from optimizing databases and messaging systems to refining our STACKIT services.

  • You collaborate closely with development teams to shorten time-to-detect intervals by enhancing our monitoring and alerting infrastructure and ensuring our services adhere to defined SLOs.
  • Your work is critical in continuously optimizing our time-to-mitigation; you achieve this by creating clear playbooks, designing dashboards for first responders, and ensuring our telemetry data (logs and metrics) is comprehensive.
  • You act as a reliability consultant to development teams, educating them on reliability patterns and helping them "shift left" to foster a shared responsibility model.
  • You design and refine development practices, including CI/CD pipelines, to support progressive delivery strategies such as Canary releases and Blue/Green deployments.
  • You proactively analyze and optimize the scalability of the Control Plane, addressing bottlenecks in distributed consensus, database throughput, and kernel-level networking.
  • You participate in a compensated on-call rotation, leading incident responses and facilitating blameless post-mortems and Root Cause Analyses.
  • You bring 3+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a specific focus on operating large-scale distributed systems in production.
  • You possess expert-level knowledge of Kubernetes Control Plane internals, including the API Server, Controller Manager, Scheduler, and etcd.
  • You demonstrate proficiency in Go and write production-grade code to build automation tools, Kubernetes Operators, or glue code that integrates disparate systems.
  • You hold deep experience with Infrastructure as Code and container infrastructure, alongside proficiency in Linux system internals (kernel tuning, memory management) and networking (TCP/IP, CNI, Load Balancers, eBPF).
  • You bring experience in operating datastores (e.g., PostgreSQL, Redis) and messaging systems (e.g., Kafka, NATS) in scalable environments.
  • You run towards fires to learn from them, you automate yourself out of a job, and you believe that hope is not a strategy.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, Kubernetes Control Plane & Cloud Reliability
Senior SRE, Kubernetes Control Plane & Cloud Reliability

Schwarz Dienstleistung KG • Germany (OH)

On-site
USD 104,000 - 150,000
STACKIT Platform Services - Observability Suite - Engineering Manager (m/f/d)
STACKIT Platform Services - Observability Suite - Engineering Manager (m/f/d)

Schwarz Dienstleistung KG • Germany (OH)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer (m/f/x)
Senior Site Reliability Engineer (m/f/x)

SysEleven GmbH • Germany (OH)

On-site
USD 100,000 - 130,000
Senior SRE: Kubernetes Reliability & GitOps Lead
Senior SRE: Kubernetes Reliability & GitOps Lead

SysEleven GmbH • Germany (OH)

On-site
USD 100,000 - 130,000
Site Reliability Engineer II
Site Reliability Engineer II

Kastle Systems • Falls Church (VA)

On-site
USD 120,000 - 160,000
Medical/Dental/Vision
401K
Tuition/Training Assistance
+2
Senior Site Reliability Engineer (80–100%)
Senior Site Reliability Engineer (80–100%)

Open Systems • Little Switzerland (OR)

On-site
USD 102,798 - 171,330
STACKIT Access Domain - Cloud Software Engineer (m/f/d)
STACKIT Access Domain - Cloud Software Engineer (m/f/d)

Schwarz Dienstleistung KG • Germany (OH)

On-site
USD 90,000 - 130,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Senior Site Reliability Engineer (m/w/d)
Senior Site Reliability Engineer (m/w/d)

Impower • Germany (OH)

Hybrid
USD 81,000 - 105,000
Flexible working hours
Real ownership in projects
Growth opportunities in a modern environment
Site Reliability Engineer
Site Reliability Engineer

FORT • United States

Hybrid
USD 150,000 - 180,000
Healthcare benefits
Flexible work environment
Large-scale cloud platform project
+1