SRE: Sovereign AI Cloud Platform Reliability

Nebul

Leiden

On-site

EUR 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebul in Leiden is seeking a Site Reliability Engineer to keep our sovereign AI cloud stable, observable and scalable. You’ll own incident response, drive automation and collaborate with Go and Python services across Kubernetes and GPU infrastructure.

Join a focused team delivering robust platform reliability, monitoring, and improve production readiness while reducing toil for senior engineers.

Qualifications

  • Must have hands-on production experience with SRE/DevOps platforms.
  • Experience with incident management and root cause analysis.
  • Strong grasp of monitoring, logging, metrics and alerting.

Responsibilities

  • Monitor and improve reliability, availability and performance of Nebul’s AI cloud platform.
  • Troubleshoot incidents across Kubernetes, GPU infrastructure, Linux, networking and platform services.
  • Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.
  • Investigate issues affecting services written in Go and Python.
  • Automate repetitive operational tasks using Go, Python or scripting.

Skills

Kubernetes
Linux troubleshooting
Go or Python
Monitoring & alerting
Incident response
Automation scripting
Cloud-native systems
Ownership & leadership

Tools

Go
Python

Job description

Nebul in Leiden is seeking a Site Reliability Engineer to keep our sovereign AI cloud stable, observable and scalable. You’ll own incident response, drive automation and collaborate with Go and Python services across Kubernetes and GPU infrastructure.

Join a focused team delivering robust platform reliability, monitoring, and improve production readiness while reducing toil for senior engineers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – AI Cloud Platform
Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden

On-site
EUR 70,000 - 110,000
Senior Network SRE: Build Reliable AI Cloud Backbone
Senior Network SRE: Build Reliable AI Cloud Backbone

ApplyMint • Netherlands

Remote
EUR 90,000 - 130,000
Competitive compensation
Career growth
Flexibility
+3
SRE Team Lead - Reliability, Automation & Scale
SRE Team Lead - Reliability, Automation & Scale

Together AI • Amsterdam

On-site
EUR 80,000 - 120,000
Senior Cloud SRE — Scalable AI Platform & Reliability
Senior Cloud SRE — Scalable AI Platform & Reliability

Mistral • Amsterdam

On-site
EUR 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Triwill Group • Netherlands

On-site
EUR 90,000 - 130,000
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral • Amsterdam

On-site
EUR 90,000 - 130,000
SRE: AI-Driven Cloud Observability & SLO Architect
SRE: AI-Driven Cloud Observability & SLO Architect

12Build • Netherlands

Hybrid
EUR 61,000 - 73,000
Startsalaris 5.500–6.500€
Lease a Bike
Thuiswerkvergoeding
+7
Senior Site Reliability Engineer — Remote (AWS, Kubernetes)
Senior Site Reliability Engineer — Remote (AWS, Kubernetes)

Jobgether • Netherlands

On-site
EUR 120,000 - 160,000
Fully remote within Europe
Technical ownership
Cross-team collaboration
Sovereign Cloud Network Automation Engineer
Sovereign Cloud Network Automation Engineer

Nebul • Leiden

On-site
EUR 70,000 - 110,000
Senior SRE: Observability, Automation & Resilience
Senior SRE: Observability, Automation & Resilience

Replit • Netherlands

Hybrid
EUR 70,000 - 100,000
Competitive Salary
401(k) Program
Health, Dental, Vision Insurance
+3