Senior site reliability engineer

Vacant

Hoofddorp

Hybrid

EUR 90,000 - 140,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

HeadFirst x Impellam seeks a Senior Site Reliability Engineer to keep our Azure-based cloud platform reliable, scalable, and resilient. You will work hands-on with Cloud, Data, and AI teams to automate operations, improve observability, and prevent incidents through engineering excellence.

The role emphasizes staying hands-on, with no people-management responsibilities, and collaborating across Cloud, Data, and AI to raise engineering standards and platform reliability.

Qualifications

  • 7+ years of hands-on experience in Site Reliability/Cloud Engineering roles.
  • Strong expertise in Microsoft Azure and cloud-native architectures.
  • Proven ability to design and operate CI/CD pipelines and IaC.
  • Deep experience with observability tools and incident response.
  • Experience with Databricks and data/AI integrations a plus.

Responsibilities

  • Improve reliability, availability, and performance of Azure-based platforms.
  • Build and enhance monitoring, logging, and alerting (Grafana, OpenTelemetry, Azure Monitor).
  • Automate tasks with Infrastructure as Code and scripting (Python, Bash).
  • Develop self-healing capabilities and automated remediation to reduce incidents.
  • Lead root cause analyses and implement long-term improvements.
  • Define and refine SLIs/SLOs and drive operational excellence.

Skills

Azure
Terraform
GitHub Actions
Grafana
OpenTelemetry
Python
Bash
CI/CD
Databricks
Incident Management
GitOps

Tools

Azure Databricks
Terraform
GitHub Actions

Job description

Senior Site Reliability Engineer

Keep our global cloud platform reliable, scalable, and resilient. Do you enjoy solving complex production challenges before they turn into incidents? Are you the kind of engineer who automates repetitive tasks, improves reliability through engineering, and believes that every outage is an opportunity to build a better system? And are you an experienced, hands-on engineer who wants to stay close to the technology rather than move into people management? We’re looking for a Senior Site Reliability Engineer to join the Global Platform Team at HeadFirst x Impellam. In this role, you’ll help build and operate the cloud platform that powers our global Workforce-as-a-Service ecosystem. You’ll work alongside Cloud Engineers, Platform Engineers, Data Engineers, and AI specialists to improve platform resilience, reduce operational overhead, and ensure that our Azure-based engineering environment remains reliable, scalable, and resilient. This is a hands-on individual contributor role, with no people-management responsibilities.

Your impact

As a Senior Site Reliability Engineer, your focus is to keep our platforms healthy, reliable, and easy to operate while continuously improving how we build, run and recover production systems.

You’ll help build the engineering foundations behind our Headless Data Architecture (HDA), which runs on Azure and Databricks, as well as the Custom Apps Infrastructure (CA) that powers integrations, internal applications, and operational workflows across our international organization.

Instead of spending your days reacting to incidents, you’ll focus on preventing them through automation, observability, and reliability engineering. You’ll reduce operational burden, improve platform resilience, and build systems that scale, recover automatically whenever possible, and provide engineering teams across Cloud, Data, and AI with the visibility they need to run production workloads with confidence.

What you will do
  • Improve the reliability, availability, and performance of our Azure platform and production environments.
  • Build and improve monitoring, logging, and alerting using Grafana, OpenTelemetry, Azure Monitor, and Log Analytics.
  • Automate operational tasks and eliminate repetitive manual work using Infrastructure as Code and scripting.
  • Design self-healing capabilities and automated remediation to reduce incidents and improve recovery times.
  • Investigate production incidents, perform root cause analyses, and implement long-term improvements.
  • Define, measure, and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Optimize platform performance, scalability, and operational efficiency.
  • Work closely with Cloud Engineers to improve platform design, resilience, security, and operational reliability.
  • Support Data and AI teams by improving the reliability of Azure Databricks environments.
  • Help shape and raise engineering standards around observability, automation, and operational excellence.
  • Continuously look for opportunities to reduce operational complexity and improve developer productivity.
  • Act as a technical point of reference for reliability engineering, helping shape technical decisions while remaining hands-on in the technology.
About the role

As part of the Global Platform Team, you’ll work alongside engineers in Cloud, Data, and AI to improve the reliability of our Azure-based platform. Using technologies such as Kubernetes, Terraform, Databricks, GitHub Actions, Grafana, and OpenTelemetry, you’ll help ensure that our global Workforce-as-a-Service ecosystem remains reliable, scalable, and resilient.

This is a hands-on individual contributor role. You will have significant technical ownership and will be expected to get into the detail when needed, from troubleshooting complex production issues to improving automation, observability and system reliability. You will act as a technical reference point for others, but you will not have people-management responsibilities.

About HeadFirst x Impellam

HeadFirst x Impellam is one of Europe’s leading providers of workforce and talent solutions. Operating across multiple countries, we are transforming into a cloud-native, AI-powered organization that connects people, technology, and data through a modern digital platform. The Global Platform Team is at the heart of that transformation, enabling engineering teams across Cloud, Data, AI, and Software Engineering to build and operate scalable solutions for the future.

What you’ll bring
  • You’re passionate about building reliable systems and solving operational challenges through engineering rather than manual intervention.
  • You enjoy understanding how distributed systems behave, thrive in cloud-native environments, and are always looking for ways to improve automation, resilience, and observability.
  • You stay calm under pressure, take ownership of problems, and enjoy collaborating with others to continuously improve the platform’s reliability.
  • You have significant experience operating production systems and still enjoy getting hands-on.
  • You are comfortable working independently, making technical decisions and going deep into complex problems.
  • While you may naturally act as a technical reference point for others, you are not looking for a people-management role.
  • 7+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps Engineer, or Cloud Engineer, with significant hands-on experience in production environments.
  • Strong hands-on experience with Microsoft Azure.
  • Experience with Infrastructure as Code using Terraform.
  • Experience building and maintaining CI/CD pipelines using GitHub Actions or Azure DevOps.
  • Experience with observability tools such as Grafana, OpenTelemetry, Azure Monitor, or Log Analytics.
  • Strong scripting skills using Python, Bash, or similar languages.
  • Experience supporting distributed cloud platforms in production.
  • Experience with incident management, root cause analysis, and post-incident improvements.
  • Familiarity with GitOps principles and modern deployment practices.
  • Experience with Azure Databricks is a strong advantage.
  • Experience with SnapLogic or similar integration platforms is a plus.
  • A strong preference for staying hands-on and solving technical problems directly, rather than moving into people-management.
Practical information

Location: Hoofddorp, North Holland, Netherlands

Working model: Hybrid, 60/40, with regular presence at our Hoofddorp office

Language: English is required. Dutch is a plus

Work authorisation: You should already be based in the Netherlands. Visa sponsorship may be available where applicable. We do not provide relocation support

Mobility budget or Lease car: No

Brand: HeadFirst Group

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

TOPdesk • Delft

Hybrid
EUR 100,000 - 130,000
Hybrid work environment
10 to Grow programme
Excellent employment conditions
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Topdesk-7 • Delft

Hybrid
EUR 59,000 - 86,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cluster - Data professionals • Den Haag

On-site
EUR 61,000 - 102,000
Hybrid work
Home office allowance
Pension scheme
+8
Senior DevOps Engineer
Senior DevOps Engineer

TKH AI Technology • Amsterdam

Hybrid
EUR 90,000 - 120,000
Hybrid work in Amsterdam
Competitive salary
Career growth opportunities
DevOps/Platform Engineer
DevOps/Platform Engineer

Wellis • Rotterdam

Hybrid
EUR 61,000 - 89,000
Salary range €5,500 – €8,000 gross per
Pension plan
Learning budget
+7
Senior Site Reliability Engineer
Senior Site Reliability Engineer

DataSnipper • Amsterdam

On-site
EUR 90,000 - 140,000
Equity
Pension
Vacation days
+7
DevOps Engineer
DevOps Engineer

Wellis Health • Rotterdam

On-site
EUR 61,000 - 89,000
Hybrid work
Pension plan
Personal budget
+3
DevOps/Platform Engineer
DevOps/Platform Engineer

Wellis Health • Rotterdam

Hybrid
EUR 61,000 - 89,000
Salary range
Pension plan
Learning budget
+6
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • Netherlands

Remote
EUR 90,000 - 130,000
Fully remote work
Lead Azure Engineer
Lead Azure Engineer

EPAM Systems Inc • Netherlands

Hybrid
EUR 110,000 - 140,000