Senior Site Reliability Engineer / Kubernetes

Jobgether

Deutschland

Remote

EUR 90.000 - 120.000

Vollzeit

Vor 8 Tagen
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobgether is seeking a Senior Site Reliability Engineer based in Germany to own and operate production Kubernetes platforms across on-premise and cloud-like environments. You will design networking, automate provisioning, and lead incident response with a focus on reliability and observability.

The role is remote-first within the EU and requires strong automation and cross-team collaboration. The position emphasizes ownership, scalability, and improving platform availability, with opportunities

Qualifikationen

  • Expert-level, hands-on experience operating Kubernetes in production environments with cluster architecture, lifecycle, networking, storage, and security.

Aufgaben

  • Operate, maintain, and optimize Linux-based infrastructure (Debian/Ubuntu) for security, stability, and performance.

Kenntnisse

Kubernetes
Linux administration
Networking
Automation
Observability
Incidents
GitOps
On-call
English
Cloud infrastructure

Tools

OpenStack
Proxmox
VMware
MAAS
PXE
cloud-init
Prometheus
Grafana
ELK
Loki
Graylog
Istio
Linkerd
Cloudflare

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer / Kubernetes based in Germany.

This is a high-impact SRE role focused on building and operating reliable, scalable cloud infrastructure across complex environments.
You will take ownership of production Kubernetes platforms spanning bare-metal, virtualized, and on-premise infrastructure.
The role combines deep systems and networking expertise with automation, observability, incident response, and reliability engineering.
You will help shape infrastructure architecture, improve availability and performance, and establish operational standards that scale with growth.
Working closely with engineering and cross-functional teams, you will turn operational challenges into robust, repeatable solutions.
The environment is highly technical, international, remote-first, and designed for engineers who value autonomy and ownership.
If you enjoy solving complex infrastructure problems and building dependable platforms from the ground up, this role offers significant scope and impact.

Accountabilities
  • Operate, maintain, and continuously improve Linux-based infrastructure, primarily across Debian and Ubuntu environments, ensuring systems remain secure, stable, and performant.
  • Deploy, manage, and scale production Kubernetes clusters across bare-metal, virtualized, and on-premise environments, taking ownership of the full cluster lifecycle including upgrades, node pools, networking, storage, and security hardening.
  • Design and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity, with a strong focus on reliability and performance.
  • Build and maintain automated infrastructure provisioning and operational workflows using Ansible, Bash, Python, Git, and GitOps practices, including PXE boot, Preseed, and cloud-init processes.
  • Develop and operate comprehensive observability platforms using technologies such as Prometheus, Grafana, Loki, ELK, and Graylog, improving monitoring quality and ensuring alerts provide actionable operational insights.
  • Lead incident response and escalation activities, investigate complex infrastructure issues, reduce recurring failures, and improve overall platform availability and latency.
  • Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
  • Establish and maintain effective on-call practices and schedules that provide reliable operational coverage across multiple time zones.
  • Create Standard Operating Procedures and repeatable operational processes for infrastructure maintenance, troubleshooting, incident response, and routine platform activities.
  • Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance activities, and data-center operations.
  • Manage virtualization and orchestration technologies including OpenStack, Proxmox, and VMware, while contributing to the broader architecture of the platform and its products.
  • Plan infrastructure capacity and resources for future initiatives, taking projected demand, scalability, and growth into account.
  • Partner with development and cross-functional teams to improve software and infrastructure quality, optimize resource utilization, and strengthen operational practices across the wider organization.
Requirements:
  • Expert-level, hands-on experience operating Kubernetes in production environments, with a strong understanding of cluster architecture, lifecycle management, networking, storage, and security.
  • Strong network engineering expertise is essential, including VLANs, L2/L3 routing, VPNs, multi-site connectivity, and the ability to design and operate complex network architectures.
  • Advanced Linux systems administration experience, particularly with Debian and Ubuntu environments.
  • Strong understanding of networking fundamentals, distributed systems, container orchestration, and modern infrastructure architecture.
  • Proven experience building and maintaining infrastructure automation using Ansible, Bash and/or Python, Git-based workflows, and GitOps practices.
  • Practical experience with observability and monitoring platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Experience with virtualization technologies including OpenStack, Proxmox, and/or VMware.
  • Experience with bare-metal provisioning and technologies such as MAAS, PXE, Preseed, and cloud-init.
  • Experience managing incidents, escalation processes, on-call rotations, and operational reliability in production environments.
  • A process-oriented mindset with the ability to create SOPs and operational procedures from the ground up.
  • Strong problem-solving and analytical skills, combined with the ability to work independently in a fast-paced, engineering-driven environment.
  • Excellent communication and collaboration skills, with the ability to work effectively with distributed, cross-functional teams.
  • Fluent English is mandatory, with the ability to communicate complex technical topics clearly.
  • Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
  • Knowledge of Cloudflare APIs, DNS automation, or tunnel configuration is advantageous.
  • Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls and network policies, or IT asset management is beneficial.
  • Experience working across multiple time zones or establishing SRE and reliability frameworks within growing organizations is a strong advantage.
Benefits:
  • 100% remote work within the EU time zone, with a preferred working range of CET ±2 hours.
  • Flexible working hours designed to support autonomy and effective collaboration across distributed teams.
  • High-impact position with significant ownership, responsibility, and freedom to shape infrastructure and reliability practices.
  • Opportunity to work with advanced Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability technologies.
  • Collaborative and international engineering environment with strong technical expertise and cross-functional exposure.
  • Opportunity to influence architecture, operational standards, automation strategies, and long-term platform scalability.
  • Fast-paced environment focused on reliability, engineering excellence, innovation, and continuous improvement.
  • Opportunity to contribute to ambitious infrastructure projects while developing deeper expertise across modern cloud and SRE technologies.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FactFinder • Berlin

Vor Ort
Confidential
Hybrid work model
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 125.000
Hybrid work model
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Berlin

Hybrid
EUR 110.000 - 150.000
Hybrid work model
Flexible work policy
AI-driven environment
Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 120.000
Competitive pay
Career growth
Flexible work
+2
Site Reliability Engineer (SRE) – Kubernetes/Platform - Berlin/Frankfurt - €110,000–120,000
Site Reliability Engineer (SRE) – Kubernetes/Platform - Berlin/Frankfurt - €110,000–120,000

Findr • Berlin

Hybrid
EUR 110.000 - 120.000
Site Reliability Engineer
Site Reliability Engineer

Roc Search • Deutschland

Vor Ort
EUR 70.000 - 110.000
Senior Site Reliability Engineer / SRE - Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE - Kubernetes & Hybrid Cloud (m/f/d)

FactFinder • Berlin

Vor Ort
EUR 90.000 - 140.000
Hybrid work model (3 office days/week)
Senior Site Reliability Engineer / SRE - Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE - Kubernetes & Hybrid Cloud (m/f/d)

Fact Finder • Berlin

Vor Ort
EUR 90.000 - 120.000
Sovereign Cloud Engineer (m/w/d)
Sovereign Cloud Engineer (m/w/d)

TMT Prüfservice GmbH & Co KG • Walldorf

Hybrid
Confidential
Site Reliability Engineer (f/m/d)
Site Reliability Engineer (f/m/d)

1&1 IONOS SE • Karlsruhe

Hybrid
EUR 70.000 - 100.000
Hybrid work
Flexible hours
Canteen subsidy (at some locations)
+5