Senior Site Reliability Engineer

Cluster - Data professionals

Den Haag

On-site

EUR 61,000 - 102,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work
Home office allowance
Pension scheme
Personal development budget
Vitality Budget for health
Travel expenses / NS card
Company bicycle lease
Senior technical role
Work on international SaaS platform
10% time for personal development
Laptop and phone provided

Job summary

Cluster - Data professionals is seeking a Senior Site Reliability Engineer to ensure the reliability of a large international SaaS environment on Microsoft Azure. You will shape SLOs, automate toil, and enhance observability while partnering with product and engineering to drive scalable, secure deployments.

You will contribute to incident response, canary deployments, and cost forecasting, with a hybrid work setup and competitive benefits.

Qualifications

  • 5+ years of experience as a Site Reliability Engineer, DevOps Engineer, Cloud Engineer or Infrastructure Engineer.
  • Hands-on experience managing production cloud environments, preferably Microsoft Azure.
  • Experience with SLOs, error budgets and reliability engineering principles.
  • Strong observability and monitoring knowledge, e.g. Grafana/Prometheus.
  • Hands-on Kubernetes experience, incl. Helm, ingress, RBAC and persistent volumes.
  • Terraform and Infrastructure as Code experience.
  • Python or another language for automation.
  • Experience with CI/CD environments.
  • Ability to own production incidents and participate in on-call rotation.
  • Strong communication and documentation skills.

Responsibilities

  • Define and manage SLOs and error budgets for critical services across the Azure environment.
  • Identify operational toil and replace repetitive manual work with automation and self-healing solutions.
  • Improve observability, including metrics, alerting and tracing.
  • Lead complex production incidents and facilitate blameless postmortems.
  • Translate lessons from incidents into structural improvements and automation.
  • Improve CI/CD and progressive delivery through safe deployments, canary releases and automated rollback.
  • Manage capacity, performance and infrastructure cost forecasting across the platform.
  • Build reliability automation using Python and Terraform.
  • Introduce AI agents and bounded automation into detection, diagnosis and remediation.
  • Work proactively with product and engineering teams to incorporate SRE principles into new solutions from the design stage.

Skills

SRE principles
SLOs and budgets
Observability
Python automation
CI/CD
Kubernetes
Terraform
Infra as code
On-call ownership
Communication skills

Tools

Grafana
Prometheus
Helm
RBAC
Ingress
PVs

Job description

As a Senior Site Reliability Engineer, you will be responsible for the reliability of a large international SaaS environment running primarily on Microsoft Azure. The platform operates across more than 10 global data centres, serves millions of end users and needs to deliver 24/7 availability backed by strict SLAs.

This is not a traditional operations or ticket-driven infrastructure role. You will approach reliability from an engineering perspective. You define and manage SLOs and error budgets, improve observability, automate repetitive work and design systems that can recover automatically when something goes wrong.

You will work closely with cloud engineers and software development teams. Instead of only becoming involved after an incident occurs, you will participate early in the development process and help engineering teams make architectural decisions that improve scalability, performance and reliability.

What will you do?
  • Define and manage SLOs and error budgets for critical services across the Azure environment
  • Identify operational toil and replace repetitive manual work with automation and self-healing solutions
  • Improve and standardise observability, including metrics, alerting and tracing
  • Lead complex production incidents and facilitate blameless postmortems
  • Translate lessons from incidents into structural improvements and automation
  • Improve CI/CD and progressive delivery through safe deployments, canary releases and automated rollback
  • Manage capacity, performance and infrastructure cost forecasting across the platform
  • Build reliability automation using technologies such as Python and Terraform
  • Introduce AI agents and bounded automation into detection, diagnosis and remediation
  • Work proactively with product and engineering teams to incorporate SRE principles into new solutions from the design stage
Who are you?

You are an experienced engineer who enjoys taking ownership of complex production environments. You don't just want to keep systems running; you want to understand why problems occur and engineer them out of the environment.Ideally, you bring:

  • Around 5+ years of experience as a Site Reliability Engineer, DevOps Engineer, Cloud Engineer or Infrastructure Engineer
  • Hands-on experience managing production cloud environments, preferably Microsoft Azure
  • Experience working with SLOs, error budgets and reliability engineering principles
  • Strong knowledge of observability and monitoring, for example Grafana, Prometheus or similar tooling
  • Hands-on Kubernetes experience, including areas such as Helm, ingress, RBAC and persistent volumes
  • Experience with Terraform and Infrastructure as Code
  • Experience using Python or another language for automation
  • Experience working with CI/CD environments
  • The confidence to take ownership during production incidents and participate in an on-call rotation
  • Strong communication skills and the ability to document technical decisions, incidents and architecture clearly
About the organization

You will join an international software company that develops service management software for organizations across sectors such as government, education, healthcare and industry.

The organization employs more than 700 people across eight international offices, while its software is used by more than 10 million users worldwide.

The working environment is characterized by limited hierarchy, significant individual responsibility and close collaboration between engineering teams. Technology and innovation are central to the organization, with continuous investment in its SaaS platform, automation and the adoption of AI within both its products and engineering processes.

What do we offer?
  • Salary up to €102.000,-
  • 32 to 40-hour working week
  • Hybrid working, including a home office allowance
  • A pension scheme with employer contribution
  • A personal development budget for conferences, certifications, and courses
  • A Vitality Budget of €50 per month for sports, gym membership, or other things that contribute to your mental and physical health
  • Travel expense reimbursement or an NS Business Card, plus the option to lease a bicycle through the company
  • A senior technical position with significant ownership
  • The opportunity to work on an international SaaS platform used by millions of users
  • 10% of your working time dedicated to personal development
  • A personal development budget equivalent to 10% of your gross annual salary
  • A modern laptop and phone
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Topdesk-7 • Delft

Hybrid
EUR 59,000 - 86,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Cluster - Data professionals • Den Haag

Hybrid
EUR 100,000 - 125,000
Hybrid working
Home office allowance
Pension scheme with employer contrib
+6
Infrastructure Engineering Manager
Infrastructure Engineering Manager

Elevation Group • Den Haag

Hybrid
EUR 96,000 - 118,000
Hybrid working
Pension scheme
Development budget
+5
Azure Cloud Engineer
Azure Cloud Engineer

Exact Software • Delft

Hybrid
EUR 70,000 - 95,000
Competitive salary package
13th month
8% holiday allowance
+6
Senior Site Reliability Engineer/Platform Engineer – Data Platforms Amsterdam, hybrid (3 days i[...]
Senior Site Reliability Engineer/Platform Engineer – Data Platforms Amsterdam, hybrid (3 days i[...]

sHR Consultancy • Netherlands

On-site
EUR 80,000 - 100,000
Permanent contract
20% guaranteed bonus in stock
12K annual allowance
+2
EC | Cloud Engineer Azure | SaaS | Utrecht | 90k
EC | Cloud Engineer Azure | SaaS | Utrecht | 90k

WKL Consultancy • Utrecht

Hybrid
EUR 54,000 - 90,000
Salary up to €90,000 gross per year
Hybrid work environment
13th month salary
+5
Senior Azure Cloud Engineer | Rotterdam | €5,400 – €6,950
Senior Azure Cloud Engineer | Rotterdam | €5,400 – €6,950

DBR Groep • Rotterdam

Hybrid
EUR 70,000 - 90,000
EC | Cloud Application Specialist | SaaS | Utrecht | 90k
EC | Cloud Application Specialist | SaaS | Utrecht | 90k

WKL Consultancy • Utrecht

Hybrid
EUR 54,000 - 90,000
13th month salary
8% holiday allowance
Flexible hybrid
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Topdesk-7 • Delft

Hybrid
EUR 56,000 - 106,000
Senior Database Engineer
Senior Database Engineer

TOPdesk • Delft

Hybrid
EUR 90,000 - 140,000