Senior Site Reliability Engineer

Jobgether

Lavamünd

Remote

EUR 159.000 - 201.000

Vollzeit

vor 24 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, versandbereit.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Remote work
Flexible hours
High ownership
Modern stack
Global team

Zusammenfassung

Jobgether in Switzerland seeks a Senior Site Reliability Engineer to own production infrastructure, focusing on reliability, scalability, and automation across Kubernetes, Linux, networking, virtualization, and on‑prem environments.

You will define SLOs/SLIs, lead incident responses, and build observability stacks with Prometheus, Grafana, ELK/Loki, and Graylog, collaborating with cross‑functional teams to reduce latency and improve availability.

Qualifikationen

  • Expert-level, hands‑on experience operating Kubernetes in production.
  • Strong network engineering expertise across VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Strong Linux systems administration skills (Debian/Ubuntu).
  • Solid understanding of networking fundamentals and complex network design.
  • Proven experience with infrastructure automation (Ansible, Bash, Python), Git-based workflows, and GitOps.
  • Experience with observability platforms (Prometheus, Grafana, ELK, Loki, Graylog).
  • Experience with virtualization (OpenStack, Proxmox, VMware).
  • Experience with MAAS and bare-metal provisioning.
  • Strong distributed systems and container orchestration knowledge.
  • Process-oriented mindset to create SOPs and operational procedures.

Aufgaben

  • Operate and improve Linux-based infrastructure (Debian/Ubuntu).
  • Deploy and scale production Kubernetes clusters across bare-metal, virtualization, and on‑prem environments.
  • Design and maintain complex networking architectures (VLANs, L2/L3 routing, VPNs).
  • Build infrastructure automation using Ansible, Bash, Python; GitOps practices (PXE, Preseed, cloud-init).
  • Maintain observability and monitoring platforms (Prometheus, Grafana, ELK, Loki, Graylog).
  • Lead incident response, escalation, and on-call rotations; reduce latency and increase availability.
  • Define and implement SLOs/SLIs across infrastructure and services.
  • Create SOPs and coordinate data‑center operations and hardware maintenance.

Kenntnisse

Kubernetes production
Linux admin
Networking
GitOps
Automation
Observability
SRE practices
Incident management
On-call
Cloud/Hybrid infra

Tools

OpenStack
Proxmox
VMware
MAAS
PXE boot
Preseed
cloud-init

Jobbeschreibung

Senior Site Reliability Engineer based in Switzerland.

This is a senior infrastructure role focused on building and operating highly reliable cloud platforms in a distributed engineering environment.

You will take ownership of production infrastructure across Kubernetes, Linux, networking, virtualization, and bare-metal environments.

The role combines deep technical expertise with automation, observability, incident management, and proactive reliability engineering.

You will help shape infrastructure architecture, improve availability and performance, and establish scalable operational practices.

Working closely with engineering and cross-functional teams, you will solve complex infrastructure challenges and optimize resource utilization.

The environment is fully remote, international, and highly collaborative, with significant autonomy and ownership.

This is an opportunity to make a direct impact on ambitious cloud infrastructure projects while working with modern technologies.

Accountabilities
  • Operate, maintain, and continuously improve Linux-based infrastructure, with a strong focus on Debian and Ubuntu environments.
  • Deploy, manage, and scale production Kubernetes clusters across bare-metal, virtualized, and on-premise environments, overseeing upgrades, node pools, networking, storage, and security hardening.
  • Design, implement, and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Build and maintain infrastructure automation using Ansible, Bash, Python, Git-based workflows, and GitOps practices, including automated provisioning through PXE boot, Preseed, and cloud-init.
  • Deploy and maintain observability and monitoring platforms such as Prometheus, Grafana, Loki, ELK, and Graylog, ensuring operational data generates actionable insights.
  • Lead incident response and escalation activities, troubleshoot complex infrastructure issues, and implement improvements that increase availability and reduce latency.
  • Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
  • Optimize alerting and monitoring pipelines while establishing effective on-call schedules to provide operational coverage across time zones.
  • Create and maintain Standard Operating Procedures for recurring infrastructure operations, maintenance, troubleshooting, and incident management.
  • Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance, and data‑center operations.
  • Manage virtualization and orchestration layers using technologies such as OpenStack, Proxmox, and VMware.
  • Contribute to the overall architecture and evolution of infrastructure products, ensuring solutions remain scalable and reliable.
  • Plan infrastructure capacity and resources for future initiatives based on projected demand and business growth.
  • Partner with development teams to improve system quality, optimize resource utilization, and strengthen engineering practices.
  • Collaborate with cross‑functional stakeholders to align infrastructure priorities with broader product and customer needs.
Requirements
  • Expert-level, hands‑on experience operating Kubernetes in production, including cluster lifecycle management, networking, storage, security, and scaling.
  • Strong network engineering expertise is essential, particularly across VLANs, L2/L3 routing, VPNs, and multi‑site connectivity.
  • Strong Linux systems administration skills, particularly with Debian and Ubuntu.
  • Solid understanding of networking fundamentals and the ability to design and operate complex network architectures.
  • Proven experience developing infrastructure automation using Ansible, Bash and/or Python, Git‑based workflows, and GitOps methodologies.
  • Practical experience with observability platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Experience working with virtualization technologies including OpenStack, Proxmox, and VMware.
  • Experience with bare‑metal provisioning and MAAS (Metal as a Service).
  • Strong understanding of distributed systems and container orchestration.
  • A process‑oriented mindset, with the ability to create SOPs and operational procedures from the ground up.
  • Experience managing production incidents, escalation processes, and on‑call rotations.
  • Ability to work independently and make sound technical decisions in a fast‑paced, engineering‑driven environment.
  • Strong communication and collaboration skills, combined with a high level of technical ownership and alignment with team values.
  • Fluent English is mandatory.
  • Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
  • Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations is advantageous.
  • Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls, and network policies is beneficial.
  • Familiarity with IT asset management or license tracking workflows is an advantage.
  • Experience working across multiple time zones and establishing SRE or reliability frameworks within growing organizations is highly valued.
Benefits
  • 100% remote work within the EU time zone, with CET ±2 hours preferred.
  • Flexible working hours designed to support autonomy and effective collaboration.
  • High‑impact position with significant ownership and the opportunity to influence infrastructure strategy.
  • Opportunity to work with a modern technology stack spanning Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability.
  • Collaborative and international engineering environment with exposure to complex infrastructure challenges.
  • Significant autonomy to shape operational processes, reliability practices, and technical solutions.
  • Opportunity to contribute to ambitious cloud infrastructure initiatives with a strong focus on reliability, automation, and continuous improvement.

We appreciate your interest and wish you the best!

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Lavamünd

Vor Ort
EUR 127.000 - 190.000
Competitive compensation
Career growth
Ownership of meaningful projects
+1
Senior Software Engineer - Reliability, Infrastructure, and Tooling
Senior Software Engineer - Reliability, Infrastructure, and Tooling

Jobgether • Lavamünd

Remote
EUR 117.000 - 261.000
Equity participation
Fully remote work
Health, dental, and vision benefits
+2
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Lavamünd

Vor Ort
EUR 127.000 - 190.000
Competitive compensation
Career growth
Learning opportunities
+2
Senior Kubernetes Platform Engineer
Senior Kubernetes Platform Engineer

Jobgether • Lavamünd

Remote
EUR 70.000 - 115.000
100% remote work
International projects
Professional development
+2
Site Reliability Engineer - eFX/ Crypto
Site Reliability Engineer - eFX/ Crypto

Swissquote • Lavamünd

Hybrid
EUR 70.000 - 110.000
Senior System Engineer (Virtual Private Cloud Team)
Senior System Engineer (Virtual Private Cloud Team)

Jobgether • Lavamünd

Vor Ort
EUR 148.000 - 211.000
Competitive compensation
International environment
Ownership in your work
Senior SRE - Kubernetes & Cloud Infra (Remote)
Senior SRE - Kubernetes & Cloud Infra (Remote)

Jobgether • Lavamünd

Vor Ort
EUR 148.000 - 201.000
100% remote work within EU time zones
Flexible working hours
High ownership and autonomy
+4
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • Lavamünd

Remote
EUR 148.000 - 211.000
Fully remote working environment
Broad technical ownership across cloud
Mentoring and leadership opportunities
+1
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • Österreich

Remote
EUR 90.000 - 130.000
Fully remote work
Platform Engineer - Deployment Team
Platform Engineer - Deployment Team

Blackshark.ai GmbH • Graz

Vor Ort
EUR 60.000 - 70.000
Flexible working arrangements
Learning opportunities
Mental well-being programs
+1