Senior Site Reliability Engineer (m/f/d)

TOPdesk

Kaiserslautern

Vor Ort

EUR 90.000 - 150.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

30 days annual vacation
Remote-friendly options
Health and wellness programs
Training and onboarding support
Employee subsidies and perks

Zusammenfassung

TOPdesk in Kaiserslautern, Germany, seeks a Senior Site Reliability Engineer to own the reliability of our Azure SaaS estate. You will set SLOs, reduce toil, and drive self-healing platforms while partnering with cloud engineering and product squads to ship reliable software.

You will lead incident response, automate repeatable work, and improve observability across multi-cloud regions. The role emphasizes canaries, progressive delivery, and blameless postmortems with a focus on cost-aware

Qualifikationen

  • Proven hands-on experience (5+ years) as a Site Reliability, DevOps, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice — you have set them, not just read about them.
  • Strong observability skills at scale — Grafana/Prometheus/other tools, including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Automation experience with Python or Terraform delivered via CI/CD.
  • Linux system administration with knowledge of configuration management tools.
  • Experience leading incidents in an on-call rotation with SLA obligations.
  • Excellent written communication for postmortems and architecture notes.

Aufgaben

  • Define and own SLOs and error budgets across the Azure estate, balancing delivery and reliability.
  • Eliminate toil and implement self-healing automation feeding reliability roadmap improvements.
  • Standardize observability across datacenters; reduce alert-to-incident ratios.
  • Lead incident response, run blameless postmortems, and derive durable fixes.
  • Harden CI/CD and progressive delivery, enabling safe canaries and automated rollbacks.
  • Model capacity and performance to keep the platform scalable across regions.
  • Apply AI-native reliability practices and bounded automation in detection and remediation.
  • Document runbooks and automation opportunities to improve legibility and reuse.

Kenntnisse

SRE/DevOps: 5+ yrs
Azure cloud at scale
Observability: Grafana/Prometheus
Kubernetes: Helm/RBAC
Python/Terraform automation
Linux administration
Incident leadership on-call
Written communication

Tools

Terraform
Puppet/Ansible
CI/CD pipelines

Jobbeschreibung

Company Description

TOPdesk builds service management software used across education, healthcare, government, and manufacturing. We are 700+ colleagues in 8 offices worldwide. Founded over 30 years ago, we serve more than 10 million users worldwide and have been helping organisations deliver better services ever since.

We are an open, collaborative organisation with little hierarchy — people own their work end to end and are trusted to make the decisions that matter. We are reinventing ITSM and ESM for the agentic era, building AI agents our customers can trust, and this role is part of that.

Job Description

About the role

Our Azure SaaS estate keeps service management running for thousands of organisations worldwide, under SLA-backed 24/7 availability. As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem — you set the SLOs, engineer out the toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.

You sit in the SaaS infrastructure function, working alongside cloud engineering and the product squads shipping to production. You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback — self-healing systems, not runbooks worked by hand.

What this is not A ticket-driven, break-fix ops role kept away from the code. This is reliability as engineering — you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reacting incident by incident.

The team We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance — and who treat reliability as a shared, measurable objective, not a firefight.

What you'll own
  • SLOs and error budgets. Define and own service-level objectives across the Azure (and potentially multi-cloud) estate, and use error budgets to steer the balance between shipping change and protecting reliability.
  • Toil elimination and self-healing automation. Identify toil, classify it, and engineer it out — feeding self-healing automation and your findings into the reliability roadmap. Standupanagent-basedsupportlayerthatownsrecurringtoilandcontinuouslyfeedsimprovementsbackintoreliability.
  • Observability consolidation. Standardise metrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads, and measurably reduce the alert-to-incident ratio from baseline.
  • Incident response and blameless postmortems. Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
  • Reliability of releases. Harden CI/CD and progressive delivery — canaries, safe rollouts, automated rollback — so change velocity and reliability rise together.
  • Capacity and performance. Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
  • AI-native reliability. Bring agents and bounded automation — with observability, approvals, containment, and rollback — into detection, diagnosis, and remediation.
  • Runbooks that get used. Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
  • Capacity and cost forecasting. Own capacity and cost planning across the multi-cloud estate, model usage and growth trends, and forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting to it.
How you approach the work
  • Automate what you repeat — if you have done it manually twice, the third time is a design problem.
  • Measure before optimising: SLOs, baselines, and dashboards before opinions.
  • Design for failure — assume things break, and make recovery automatic and observable.
  • Consultative, not gatekeeping: you pair with product engineering teams and transfer knowledge as you go.
  • Treat cost and reliability as joint objectives, not a forced trade-off.
  • Pro-active collaboration with product teams. You are involved in the early phases of product development, including design to help the teams make optimal choices and timely introduce appropriate SRE practices.
Technical environment
  • Scale: 10+ global datacenters; SLA-backed, 24/7 multi-tenant SaaS serving millions of end users.
  • Cloud: Azure across all production regions, with a mature landing-zone and networking architecture.
  • Compute: Kubernetes / Azure AKS alongside traditional VM infrastructure, all managed as code.
  • Infrastructure as code: Terraform via CI/CD and GitOps workflows; configuration management with Puppet and Ansible across Linux and Windows.
  • Observability: metrics, alerting, and tracing across cloud-native and self-managed layers (e.g. Grafana, Prometheus, VictoriaMetrics, Influx).
  • Automation: Python and automation tooling — and we expect you to take the reliability stack to the next level, not just operate today's.
  • Legacy: Java, MS SQL, heritage architecture — being decomposed. The SRE role is not responsible for the Java application code.
  • How we build: Claude Code as our primary AI-native SDLC tool; subagents and multi-agent workflows; MCP tool integrations; shared prompt, agent, and eval libraries.
Success in your first year
  • SLOs and error budgets are defined for the estate's critical services and actively used to steer delivery decisions.
  • The alert-to-incident ratio is measurably down, and runbooks you wrote are used by the on-call shift without escalation.
  • Toil you identified is automated — or has a credible, documented roadmap to be — and self-healing covers at least one high-frequency failure mode.
  • Postmortems produce durable fixes, not repeat incidents; recurring-incident rate is trending down against a documented baseline.
  • Product squads consult you during design, not only after incidents.
Qualifications
Required
  • Proven hands-on experience (5+ years) as a Site Reliability, DevOps, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice — you have set them, not just read about them.
  • Strong observability skills at scale — Grafana, Prometheus, VictoriaMetrics, or equivalent — including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Coding for automation (Python or equivalent) and Terraform delivered via CI/CD.
  • Linux system administration — you understand what Puppet or Ansible is doing, not just whether it ran green.
  • Comfortable leading incidents in an on-call rotation with real SLA obligations, and the maturity to know when to elevate.
  • Strong written communication — your postmortems, runbooks, and architecture notes are unambiguous.
Nice to have
  • Experience with progressive delivery — canaries, feature flags, automated rollback.
  • Experience working within or migrating toward an Azure Cloud Adoption Framework or enterprise landing-zone structure.
  • Current, personal practice of AI-native software delivery (Claude Code or equivalent).
  • Experience with EU data residency / sovereign cloud requirements.
Additional Information
What's in it for you
  • Permanent employment contract and 30 days of annual vacation
  • Pleasant working atmosphere with flat hierarchies
  • Open working atmosphere in an international environment
  • Flexible working hours within a modern working environment
  • Possibility to work remotely
  • Well-founded onboarding by a buddy
  • Time for individual training opportunities to further develop your personal strengths
  • Joint employee events and team building measures
  • Employee subsidy for gym membership
  • Company health measures such as health days or fresh fruit
  • Free drinks (coffee, tea, water)
  • Gifts on special occasions, e.g. anniversary
  • Monthly tax-free payment in the form of a Mastercard
  • Possibility of time off (sabbatical)
  • Quality time together (table football, table tennis table, massage chair)
  • Corporate benefits
  • Vacation bonus

We welcome applications from all interested parties, regardless of their ethnic and social background, age, religion, gender, disability, sexual orientation, or identity.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (m/f/d) at TOPdesk
Senior Site Reliability Engineer (m/f/d) at TOPdesk

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 120.000
Possibility to work remote
Flexible working hours
Senior Site Reliability Engineer (m/w/d)
Senior Site Reliability Engineer (m/w/d)

Impower • München

Hybrid
EUR 70.000 - 90.000
Flexible hours
Ownership in projects
Diverse team culture
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 140.000
Deutschlandticket
Vacation days
Health insurance
+7
Site Reliability Engineer (m/f/d)
Site Reliability Engineer (m/f/d)

gridscale • Köln

Vor Ort
EUR 60.000 - 90.000
32 vacation days
Flexible working hours
AI-augmented engineering tools
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

B Capital • Deutschland

Remote
EUR 46.000 - 105.000
Work from anywhere
Flexible paid time off
Mental health support services
+3
Senior Software Engineer
Senior Software Engineer

TOPdesk • Kaiserslautern

Hybrid
EUR 70.000 - 100.000
Hybrid working environment
10 to Grow personal development
Open, informal culture
+1
Engineering Manager - Site Reliability & Observability (x/f/m)
Engineering Manager - Site Reliability & Observability (x/f/m)

United States Digital Space LLC • Berlin

Hybrid
EUR 120.000 - 180.000
Deutschlandticket (Germany-wide public
health insurance
Pension scheme (bAV)
+4
Senior Database Engineer
Senior Database Engineer

TOPdesk • Kaiserslautern

Hybrid
EUR 90.000 - 120.000
Hybrid work environment
Excellent employment conditions
Engineering Manager (m/f/d)
Engineering Manager (m/f/d)

JobCubby • Kaiserslautern

Hybrid
EUR 84.000 - 95.000
Remote work up to 50%
30 days annual leave
Monthly tax-free benefits Mastercard
+1
Site Reliability Engineering Lead (f/m/d)
Site Reliability Engineering Lead (f/m/d)

United States Digital Space LLC • Bochum

Vor Ort
EUR 120.000 - 180.000
Pension plan
Paid time off
Transport reimbursement
+2