Senior Site Reliability Engineer

Topdesk-7

Delft

Hybrid

EUR 59,000 - 86,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TOPdesk is seeking a Senior Site Reliability Engineer to own the reliability of our Azure SaaS estate. You will define SLOs, reduce toil with automation, and improve platform resilience while balancing delivery speed.

In this role you will collaborate with cloud engineering and product squads, driving AI-native reliability practices, self-healing systems, and observability across data centers and regions. Hybrid work in the Netherlands is supported.

Qualifications

  • 5+ years hands-on experience as Site Reliability/DevOps/Infrastructure Engineer in a production cloud environment (Azure).
  • Fluent with SLOs, error budgets and reliability engineering practices.
  • Strong observability skills with Grafana/Prometheus and incident response capabilities.

Responsibilities

  • Define and own SLOs and error budgets across the Azure multi-cloud estate.
  • Eliminate toil through self-healing automation and automation candidates.
  • Harden CI/CD and progressive delivery to increase change velocity while protecting reliability.
  • Lead incident response with blameless postmortems and durable fixes.
  • Forecast capacity and cost across the multi-cloud estate to guide spend and scaling.

Skills

SRE experience 5+ years
Kubernetes
Python automation
Linux administration
Observability (Grafana/Prometheus)

Tools

Terraform
CI/CD tooling

Job description

  • Compensation: EUR 5,250 - EUR 7,750 - monthly
Company Description

TOPdesk builds service management software used across education, healthcare, government, and manufacturing. We are 700+ colleagues in 8 offices worldwide. Founded over 30 years ago, we serve more than10 million usersworldwide and have been helpingorganisationsdeliver better services ever since.

We are an open, collaborativeorganisationwith little hierarchy — people own their work end to end and are trusted to make the decisions that matter. We are reinventing ITSM and ESM for the agentic era, building AI agents our customers can trust, and this role is part of that.

Job Description

About the role

Our Azure SaaS estate keeps service management running for thousandsof organisationsworldwide, under SLA-backed 24/7 availability. As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem — you set the SLOs,engineer outthe toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.

You sit in theSaaSinfrastructure function, working alongside cloud engineering and the product squadsshipping to production. You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback — self-healing systems, not runbooks worked by hand.

What this isnotAticket-driven, break-fix ops role kept away from the code. This is reliability as engineering — you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reactingincidentby incident.

The team

We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance — and who treat reliability as a shared, measurableobjective, not a firefight.

Whatyou'llown

  • SLOs and error budgets.Define and own service-levelobjectivesacross the Azure(and potentiallymulti-cloud)estate, anduse error budgets to steer the balance between shipping change andprotectingreliability.
  • Toil elimination and self-healing automation.Identifytoil, classify it, and engineer it out — feeding self-healing automation and your findings into the reliability roadmap.Standupanagent-basedsupportlayerthatownsrecurringtoilandcontinuouslyfeedsimprovementsbackintoreliability.
  • Observability consolidation.Standardisemetrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads,and measurablyreducethe alert-to-incident ratio from baseline.
  • Incident response and blameless postmortems.Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
  • Reliability of releases.Harden CI/CD and progressive delivery — canaries, safe rollouts, automated rollback — so change velocity and reliability rise together.
  • Capacity and performance.Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
  • AI-native reliability.Bring agents andboundedautomation — with observability, approvals, containment, and rollback — into detection, diagnosis, and remediation.
  • Runbooks that get used.Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
  • Capacity and cost forecasting.Own capacity and cost planning across the multi-cloud estate, modelusageand growth trends, and forecastshort and long terminfrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting toit.

How you approach the work

  • Automate what you repeat — if you have done it manually twice, the third time is a design problem.
  • Measure beforeoptimising: SLOs, baselines, and dashboards before opinions.
  • Design for failure — assume thingsbreak, andmake recovery automatic and observable.
  • Consultative, notgatekeeping:you pair with product engineering teams and transfer knowledge as you go.
  • Treat cost and reliability as jointobjectives, not a forced trade-off.
  • Pro‑activecollaboration withproductteams.You are involved in theearly phases ofproductdevelopment, includingdesignto help theteamsmakeoptimalchoicesandtimelyintroduceappropriateSREpractices.

Technical environment

  • Scale:10+ global datacenters; SLA-backed, 24/7 multi-tenant SaaS serving millions of end users.
  • Cloud:Azure across all production regions, with a maturelanding-zoneand networking architecture.
  • Compute:Kubernetes / Azure AKS alongside traditional VM infrastructure, all managed as code.
  • Infrastructure as code:Terraform via CI/CD andGitOpsworkflows; configuration management with Puppet and Ansible across Linux and Windows.
  • Observability:metrics, alerting, and tracing across cloud-native and self-managed layers (e.g.Grafana, Prometheus,VictoriaMetrics, Influx).
  • Automation:Python and automation tooling — and we expect you to take the reliability stack to the next level, not justoperatetoday's.
  • Legacy:Java, MS SQL, heritage architecture —being decomposed.The SRE role is notresponsible for the Java application code.
  • How we build:Claude Code as our primary AI-native SDLC tool; subagents and multi-agent workflows; MCP tool integrations; shared prompt, agent, and eval libraries.

Success in your first year

  • SLOs and error budgets are defined for the estate's critical services and actively used to steer delivery decisions.
  • The alert-to-incident ratio is measurably down, and runbooks you wrote are used by the on-call shift without escalation.
  • Toil youidentifiedis automated — or has a credible, documented roadmap to be — and self-healing covers at least one high-frequency failure mode.
  • Postmortems produce durable fixes, not repeat incidents; recurring-incident rate istrending downagainst a documented baseline.
  • Product squads consult you during design, not only after incidents.
Qualifications

Required

  • Proven hands-on experience(5+ years)as a Site Reliability, DevOps, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice — you have set them, not just read about them.
  • Strong observability skills at scale — Grafana, Prometheus,VictoriaMetrics, or equivalent — including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Coding for automation (Python or equivalent) and Terraform delivered via CI/CD.
  • Linux system administration — you understand what Puppet or Ansible is doing, not just whether it ran green.
  • Comfortable leading incidents in an on-call rotation with real SLA obligations, and the maturity to know when to elevate.
  • Strong written communication — your postmortems, runbooks, and architecture notes are unambiguous.

Nice to have

  • Experience with progressive delivery — canaries, feature flags, automated rollback.
  • Experience working within or migrating toward an Azure Cloud Adoption Framework or enterprise landing-zone structure.
  • Current, personal practice ofAI-nativesoftware delivery (Claude Code or equivalent).
  • Experience with EU data residency / sovereign cloud requirements.
Additional Information

What'sin it for you

A strong focus on personal development — including, in the Netherlands, our “10 to Grow”programme: 10% of your time and budget for your own growth.

A hybrid working environment built on freedom, trust, and responsibility.

An open, informal, and supportive culture, with collaboration across national borders.

Excellent employment conditions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cluster - Data professionals • Den Haag

Hybrid
EUR 61,000 - 102,000
Hybrid work
Home office allowance
Pension scheme
+8
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Topdesk-7 • Delft

Hybrid
EUR 56,000 - 106,000
Engineering Manager
Engineering Manager

Topdesk-7 • Delft

On-site
EUR 78,000 - 92,000
10 to Grow programme
Hybrid work environment
Excellent employment conditions
Principal Cloud Engineer (Azure)
Principal Cloud Engineer (Azure)

TOPdesk • Tilburg

Hybrid
EUR 120,000 - 170,000
Hybrid work
10 to Grow
Open culture
+1
Engineering Manager
Engineering Manager

JobCubby • Netherlands

Hybrid
EUR 78,000 - 92,000
10 to Grow programme
Hybrid working environment
Open, collaborative culture
Senior Software Engineer
Senior Software Engineer

TOPdesk • Delft

Hybrid
EUR 90,000 - 120,000
Hybrid working environment
Personal growth program
Cross-border collaboration
Senior Agentic Quality Engineer
Senior Agentic Quality Engineer

Topdesk-7 • Delft

On-site
EUR 59,000 - 73,000
Hybrid working environment
Growth & learning program (10 to Grow)
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Cluster - Data professionals • Den Haag

Hybrid
EUR 100,000 - 125,000
Hybrid working
Home office allowance
Pension scheme with employer contrib
+6
Senior Infrastructure Engineer
Senior Infrastructure Engineer

TOPdesk • Delft

On-site
EUR 90,000 - 140,000
Senior Database Engineer
Senior Database Engineer

TOPdesk • Delft

Hybrid
EUR 90,000 - 140,000