SRE Monitoring & Observability

Nebul

Leiden

Hybrid

EUR 90,000 - 150,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Mobility budget
Performance bonus
Equity participation
Hybrid work around Leiden

Job summary

Nebul is building Europe’s Intelligent Cloud and seeks a Senior Monitoring & Observability Engineer to design, deploy and operate the observability foundation of our cloud platform.

You will work with Prometheus, OpenTelemetry and the Grafana stack (Grafana, Loki, Tempo, Mimir) across multi-datacentre environments, building SLIs/SLOs and scalable monitoring for infrastructure, apps, databases and Kubernetes. This hands-on role emphasizes reliable telemetry and secure, scalable operations.

Qualifications

  • Experience in monitoring, observability, DevOps, SRE or platform engineering.
  • Hands-on with Prometheus and Grafana.
  • Experience monitoring multi-tenant Kubernetes and cloud environments.
  • Strong Linux and Kubernetes knowledge.
  • Scripting with Bash and Python.
  • Experience defining alerting, SLIs and SLOs.

Responsibilities

  • Design, deploy and maintain Nebul’s monitoring and observability platform.
  • Build scalable observability solutions for multi-datacentre and multi-tenant environments.
  • Monitor cloud infrastructure, applications, databases, Kubernetes clusters, storage and network components.
  • Develop and maintain Grafana dashboards that provide actionable operational insight.
  • Design effective alerting strategies that minimise noise and accelerate incident response.
  • Define and implement service-level indicators and service-level objectives.
  • Build centralized logging and distributed tracing capabilities.
  • Implement and manage metrics, logs and traces using Prometheus, OpenTelemetry, Loki, Tempo and Mimir.
  • Optimize telemetry collection, retention, querying and storage performance.
  • Troubleshoot complex monitoring and production issues.
  • Ensure the observability platform is highly available, secure and scalable.
  • Automate the deployment and configuration of monitoring components.
  • Improve observability standards, documentation and operational practices.
  • Collaborate with cloud, infrastructure, Kubernetes, networking, AI and application teams.
  • Help engineering teams instrument their services and make effective use of telemetry.

Skills

Prometheus
Grafana
OpenTelemetry
Kubernetes
Linux
Docker
Helm
Argo CD
Ansible
Terraform
Python
Bash
GitOps
SLIs/SLOs
High availability

Tools

Grafana
Loki
Tempo
Mimir
Prometheus
Docker
Helm
Argo CD
Ansible
Terraform
OpenTelemetry

Job description

Nebul is building Europe’s Intelligent Cloud: a sovereign cloud and AI infrastructure platform designed for organisations that require performance, control, security and data sovereignty.

We develop and operate advanced cloud, Kubernetes and GPU infrastructure across multiple European data centres. As our platform grows, complete visibility across our infrastructure, applications and services becomes increasingly critical.

We are therefore looking for a Senior Monitoring & Observability Engineer to design and operate the observability foundation of our cloud platform.

Your Role

As a Senior Monitoring & Observability Engineer, you will design, deploy and maintain monitoring and observability solutions across Nebul’s multi-datacentre and multi-tenant environment.

You will work with modern observability technologies, including Prometheus, OpenTelemetry and the Grafana stack—Grafana, Loki, Tempo and Mimir. Your work will cover infrastructure, applications, databases, Kubernetes environments, storage and network components.

This is a hands‑on engineering role for someone who understands that observability extends beyond dashboards. You will help Nebul establish meaningful service‑level indicators, reliable alerting and the telemetry required to operate complex cloud and AI infrastructure at scale.

What You Will Do
  • Design, deploy and maintain Nebul’s monitoring and observability platform
  • Build scalable observability solutions for multi‑datacentre and multi‑tenant environments
  • Monitor cloud infrastructure, applications, databases, Kubernetes clusters, storage and network components
  • Develop and maintain Grafana dashboards that provide actionable operational insight
  • Design effective alerting strategies that minimise noise and accelerate incident response
  • Define and implement service‑level indicators and service‑level objectives
  • Build centralized logging and distributed tracing capabilities
  • Implement and manage metrics, logs and traces using Prometheus, OpenTelemetry, Loki, Tempo and Mimir
  • Optimize telemetry collection, retention, querying and storage performance
  • Troubleshoot complex monitoring and production issues
  • Ensure the observability platform is highly available, secure and scalable
  • Automate the deployment and configuration of monitoring components
  • Improve observability standards, documentation and operational practices
  • Collaborate with cloud, infrastructure, Kubernetes, networking, AI and application teams
  • Help engineering teams instrument their services and make effective use of telemetry
What We Are Looking For
  • Significant professional experience in monitoring, observability, DevOps, SRE or platform engineering
  • Hands‑on experience with Prometheus and Grafana
  • Strong experience with the Grafana observability stack, including Loki, Tempo and Mimir
  • Experience implementing or operating OpenTelemetry‑based solutions
  • Strong understanding of metrics, logs, traces and their relationship within distributed systems
  • Experience monitoring multi‑tenant Kubernetes and cloud environments
  • Strong Linux and Kubernetes knowledge
  • Hands‑on experience with Docker and Helm
  • Experience with GitOps and continuous deployment tools such as Argo CD
  • Experience with automation and Infrastructure as Code using Ansible and Terraform
  • Scripting experience with Bash and Python
  • Experience defining and implementing alerting, SLIs and SLOs
  • Understanding of high availability, security and access control within observability platforms
  • Strong troubleshooting skills across infrastructure and application layers
  • The ability to take ownership of technically complex production environments
Nice to Have
  • Experience with large‑scale or high‑cardinality telemetry environments
  • Knowledge of Kubernetes operators and custom resources
  • Experience monitoring OpenStack or private‑cloud environments
  • Understanding of network monitoring and telemetry
  • Experience monitoring databases, storage platforms and GPU infrastructure
  • Knowledge of SRE practices, incident management and capacity planning
  • Experience with Cortex or migration from Cortex to Mimir
  • Experience supporting AI, machine‑learning or high‑performance computing workloads
  • Contributions to open‑source observability projects
What We Offer
  • A key role in building Europe’s sovereign cloud and AI infrastructure
  • The opportunity to shape Nebul’s monitoring and observability architecture
  • Significant ownership and influence over technical and operational decisions
  • A highly experienced and ambitious engineering environment
  • Direct collaboration with Nebul’s engineering leadership and platform teams
  • Competitive compensation
  • A mobility budget
  • A performance bonus and equity participation
  • A hybrid working environment based around our office in Leiden
  • The opportunity to grow with a rapidly scaling European technology company
Why Nebul

At Nebul, you will build the observability foundation behind Europe’s Intelligent Cloud.

You will have the opportunity to define how we monitor a growing sovereign cloud platform across multiple data centres, Kubernetes environments and AI workloads. Your work will determine how quickly we identify problems, understand system behaviour and continuously improve the performance and reliability of our services.

You will work closely with Nebul’s engineering leadership, infrastructure specialists, platform engineers and application teams. Your decisions will shape our observability architecture, reliability standards and operational culture.

This is an opportunity to build—not simply maintain—one of the core technical capabilities of an ambitious European technology company.

Ready to Build the Observability Foundation Behind Europe’s Intelligent Cloud?

Apply through Frank Poll and help Nebul create the monitoring and observability platform needed to operate secure, high-performance cloud and AI infrastructure at European scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Observability Engineer
Senior Cloud Observability Engineer

Nebul • Leiden

Hybrid
EUR 90,000 - 150,000
Mobility budget
Performance bonus
Equity participation
+1
Technical Account Manager
Technical Account Manager

Nebul • Leiden

Hybrid
EUR 90,000 - 135,000
Mobility budget
Equity participation
Hybrid Leiden office
+1
Go Development Team Lead
Go Development Team Lead

Nebul • Leiden

On-site
EUR 70,000 - 90,000
Site Reliability Engineer – AI Cloud Platform
Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden

On-site
EUR 90,000 - 120,000
Virtualization Platform Engineer
Virtualization Platform Engineer

Nebul • Leiden

Hybrid
EUR 90,000 - 130,000
Hybrid work in Leiden office
Performance bonus
Equity participation
+2
Engineering Team Lead
Engineering Team Lead

Nebul • Leiden

On-site
EUR 110,000 - 150,000
Platform Engineer
Platform Engineer

Nebul • Leiden

On-site
EUR 60,000 - 100,000
AI Observability Engineer
AI Observability Engineer

Nebius • Amsterdam

On-site
EUR 90,000 - 125,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Tech Lead – AI Backend
Tech Lead – AI Backend

Nebul • Leiden

On-site
EUR 120,000 - 180,000
Infrastructure Engineer
Infrastructure Engineer

Nebul • Leiden

On-site
EUR 55,000 - 75,000