Senior Site Reliability Engineer

Jobgether

Deutschland

Remote

EUR 90.000 - 130.000

Vollzeit

vor 21 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine vollständige Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, fertig zum Versenden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

100% remote work within EU time zone (
Flexible working hours
High‑impact ownership
Modern tech stack with Kubernetes & AI
Exposure to complex infra challenges
Autonomy to shape processes
Reliability-focused initiatives

Zusammenfassung

Jobgether is seeking a Senior Site Reliability Engineer in Germany to own production infrastructure across Kubernetes, Linux, networking, and virtualization. You will automate, observe, and improve reliability in a fully remote, international setting with strong ownership.

Responsibilities include operating Debian/Ubuntu based systems, deploying scalable Kubernetes clusters, building robust networking, and leading incident response.

Qualifikationen

  • Expert‑level Kubernetes production experience with lifecycle management, networking, storage, and security.
  • Strong Linux administration skills with Debian/Ubuntu focus.
  • Proven ability to design and operate complex network architectures.
  • Automation experience with Ansible, Bash and Python; Git/GitOps workflows.
  • Practical observability experience using Prometheus, Grafana, ELK/Loki/Graylog.
  • Experience with virtualization: OpenStack, Proxmox, VMware; MAAS knowledge a plus.
  • Fluent English and strong cross‑functional collaboration; on‑call readiness.

Aufgaben

  • Operate and continuously improve Linux‑based infrastructure across multiple environments.
  • Deploy, manage, and scale production Kubernetes clusters across diverse platforms.
  • Design and maintain complex networking architectures and security hardening.
  • Develop infrastructure automation and Git‑based workflows; automate provisioning.
  • Deploy and maintain observability platforms and derive actionable insights.
  • Lead incident response, establish SLIs/SLOs, and optimize availability and latency.
  • Create SOPs for operations and incident management; coordinate time‑zone coverage.
  • Collaborate with development teams to improve system quality and reliability.

Kenntnisse

Kubernetes
Linux administration
Networking
Automation
Observability
Incident management
GitOps
SRE practices
English fluency

Tools

OpenStack
Proxmox
VMware
MAAS
Ansible
Bash
Python
Git

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Germany.

This is a senior infrastructure role focused on building and operating highly reliable cloud platforms in a distributed engineering environment.

You will take ownership of production infrastructure across Kubernetes, Linux, networking, virtualization, and bare‑metal environments.

The role combines deep technical expertise with automation, observability, incident management, and proactive reliability engineering.

You will help shape infrastructure architecture, improve availability and performance, and establish scalable operational practices.

Working closely with engineering and cross‑functional teams, you will solve complex infrastructure challenges and optimize resource utilization.

The environment is fully remote, international, and highly collaborative, with significant autonomy and ownership.

This is an opportunity to make a direct impact on ambitious cloud infrastructure projects while working with modern technologies.

Accountabilities
  • Operate, maintain, and continuously improve Linux‑based infrastructure, with a strong focus on Debian and Ubuntu environments.
  • Deploy, manage, and scale production Kubernetes clusters across bare‑metal, virtualized, and on‑premise environments, overseeing upgrades, node pools, networking, storage, and security hardening.
  • Design, implement, and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi‑site connectivity.
  • Build and maintain infrastructure automation using Ansible, Bash, Python, Git‑based workflows, and GitOps practices, including automated provisioning through PXE boot, Preseed, and cloud‑init.
  • Deploy and maintain observability and monitoring platforms such as Prometheus, Grafana, Loki, ELK, and Graylog, ensuring operational data generates actionable insights.
  • Lead incident response and escalation activities, troubleshoot complex infrastructure issues, and implement improvements that increase availability and reduce latency.
  • Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
  • Optimize alerting and monitoring pipelines while establishing effective on‑call schedules to provide operational coverage across time zones.
  • Create and maintain Standard Operating Procedures for recurring infrastructure operations, maintenance, troubleshooting, and incident management.
  • Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance, and data‑center operations.
  • Manage virtualization and orchestration layers using technologies such as OpenStack, Proxmox, and VMware.
  • Contribute to the overall architecture and evolution of infrastructure products, ensuring solutions remain scalable and reliable.
  • Plan infrastructure capacity and resources for future initiatives based on projected demand and business growth.
  • Partner with development teams to improve system quality, optimize resource utilization, and strengthen engineering practices.
  • Collaborate with cross‑functional stakeholders to align infrastructure priorities with broader product and customer needs.
Requirements
  • Expert‑level, hands‑on experience operating Kubernetes in production, including cluster lifecycle management, networking, storage, security, and scaling.
  • Strong network engineering expertise is essential, particularly across VLANs, L2/L3 routing, VPNs, and multi‑site connectivity.
  • Strong Linux systems administration skills, particularly with Debian and Ubuntu.
  • Solid understanding of networking fundamentals and the ability to design and operate complex network architectures.
  • Proven experience developing infrastructure automation using Ansible, Bash and/or Python, Git‑based workflows, and GitOps methodologies.
  • Practical experience with observability platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Experience working with virtualization technologies including OpenStack, Proxmox, and VMware.
  • Experience with bare‑metal provisioning and MAAS (Metal as a Service).
  • Strong understanding of distributed systems and container orchestration.
  • A process‑oriented mindset, with the ability to create SOPs and operational procedures from the ground up.
  • Experience managing production incidents, escalation processes, and on‑call rotations.
  • Ability to work independently and make sound technical decisions in a fast‑paced, engineering‑driven environment.
  • Strong communication and collaboration skills, combined with a high level of technical ownership and alignment with team values.
  • Fluent English is mandatory.
  • Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
  • Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations is advantageous.
  • Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls, and network policies is beneficial.
  • Familiarity with IT asset management or license tracking workflows is an advantage.
  • Experience working across multiple time zones and establishing SRE or reliability frameworks within growing organizations is highly valued.
Benefits
  • 100% remote work within the EU time zone, with CET ±2 hours preferred.
  • Flexible working hours designed to support autonomy and effective collaboration.
  • High‑impact position with significant ownership and the opportunity to influence infrastructure strategy.
  • Opportunity to work with a modern technology stack spanning Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability.
  • Collaborative and international engineering environment with exposure to complex infrastructure challenges.
  • Significant autonomy to shape operational processes, reliability practices, and technical solutions.
  • Opportunity to contribute to ambitious cloud infrastructure initiatives with a strong focus on reliability, automation, and continuous improvement.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer / Kubernetes
Senior Site Reliability Engineer / Kubernetes

Jobgether • Deutschland

Remote
EUR 90.000 - 120.000
Senior Software Engineer - Reliability, Infrastructure, and Tooling
Senior Software Engineer - Reliability, Infrastructure, and Tooling

Jobgether • Deutschland

Remote
EUR 117.000 - 261.000
Fully remote work
Equity participation
Health, dental, and vision benefits
+5
Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 120.000
Competitive pay
Career growth
Flexible work
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Berlin

Hybrid
EUR 110.000 - 150.000
Hybrid work model
Flexible work policy
AI-driven environment
Site Reliability Engineer -Openstack (m/f/d)
Site Reliability Engineer -Openstack (m/f/d)

gridscale GmbH • Köln

Hybrid
EUR 90.000 - 130.000
32 vacation days
Home office options
Pension plan
+2
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FactFinder • Berlin

Vor Ort
Confidential
Hybrid work model
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Ageras Danmark • Berlin

Vor Ort
EUR 70.000 - 90.000
Impactful role
Growth opportunities
Collaborative team
+1
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 125.000
Hybrid work model
Sovereign Cloud Engineer (m/w/d)
Sovereign Cloud Engineer (m/w/d)

TMT Prüfservice GmbH & Co KG • Walldorf

Hybrid
Confidential