Site Reliability Engineer – AI Cloud Platform

Nebul

Leiden

On-site

EUR 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebul in Leiden is seeking a Site Reliability Engineer to keep our sovereign AI cloud stable, observable and scalable. You’ll own incident response, drive automation and collaborate with Go and Python services across Kubernetes and GPU infrastructure.

Join a focused team delivering robust platform reliability, monitoring, and improve production readiness while reducing toil for senior engineers.

Qualifications

  • Must have hands-on production experience with SRE/DevOps platforms.
  • Experience with incident management and root cause analysis.
  • Strong grasp of monitoring, logging, metrics and alerting.

Responsibilities

  • Monitor and improve reliability, availability and performance of Nebul’s AI cloud platform.
  • Troubleshoot incidents across Kubernetes, GPU infrastructure, Linux, networking and platform services.
  • Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.
  • Investigate issues affecting services written in Go and Python.
  • Automate repetitive operational tasks using Go, Python or scripting.

Skills

Kubernetes
Linux troubleshooting
Go or Python
Monitoring & alerting
Incident response
Automation scripting
Cloud-native systems
Ownership & leadership

Tools

Go
Python

Job description

At Nebul, we’re building Europe’s sovereign AI cloud — trusted, secure, and purpose-built for the next generation of intelligent infrastructure.

Our platform combines Kubernetes, NVIDIA GPU infrastructure, cloud-native services and software written in Go and Python. As the platform grows, we need to increase our internal reliability capability and reduce the number of operational issues that depend on a small group of senior engineers.

What You’ll Be Doing

As a Site Reliability Engineer, you’ll work approximately 70% on site reliability and operational troubleshooting and 30% on broader DevOps and platform engineering.

Your primary responsibility will be to keep Nebul’s AI cloud platform stable, observable and operationally scalable. You’ll take ownership of incidents, investigate complex issues and ensure that senior engineers are not pulled into every troubleshooting session.

You’ll work across Kubernetes, NVIDIA GPU infrastructure, services written in Go and Python, networking and cloud-native platform components.

The role is not only about responding to incidents. You’ll identify recurring problems, automate operational work and improve the platform so that the same issues do not continue to return.

Key Responsibilities

  • Monitor and improve the reliability, availability and performance of Nebul’s AI cloud platform.
  • Troubleshoot incidents across Kubernetes, NVIDIA GPU infrastructure, Linux, networking and platform services.
  • Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.
  • Investigate issues affecting services written in Go and Python.
  • Perform root-cause analysis and translate findings into structural platform improvements.
  • Build and improve monitoring, metrics, logging, tracing and alerting.
  • Ensure alerts are relevant, actionable and connected to clear operational procedures.
  • Create runbooks, escalation procedures and troubleshooting documentation.
  • Automate repetitive operational tasks using Go, Python or scripting.
  • Support Kubernetes cluster operations, deployments, upgrades and platform changes.
  • Improve the production readiness of new services and infrastructure components.
  • Work with engineering teams to improve resilience, observability and failure handling.
  • Identify reliability risks before they result in platform or customer impact.
  • Support operational improvements across GPU workloads and NVIDIA-based infrastructure.
  • Reduce the operational dependency on senior platform engineers.
  • Contribute to broader DevOps work when additional capacity is needed within the team.

What Your Day Will Not Look Like

  • Acting as a first-line support engineer who only closes tickets.
  • Escalating every complex issue another senior engineer.
  • Spending all your time manually operating Kubernetes.
  • Resolving incidents without addressing their underlying causes.
  • Building isolated automation that is not integrated into the platform.

What You Bring

  • Strong experience as a Site Reliability Engineer, DevOps Engineer or Platform Engineer.
  • Hands-on production experience with Kubernetes.
  • Strong Linux and infrastructure troubleshooting skills.
  • Experience investigating issues across applications, infrastructure, networking and cloud platforms.
  • Experience with services written in Go or Python.
  • Practical experience with monitoring, logging, metrics and alerting.
  • Experience responding to production incidents and performing root-cause analysis.
  • The ability to independently lead complex troubleshooting sessions.
  • Experience automating operational work using Python, Go or scripting.
  • A solid understanding of cloud-native and distributed systems.
  • A calm, analytical and structured approach to incidents.
  • An ownership mindset and the ability to move from reactive troubleshooting to lasting improvements.

Bonus Points If You Have

  • Experience with NVIDIA GPU infrastructure.
  • Experience supporting AI, machine-learning or high-performance computing workloads.
  • Knowledge of Kubernetes GPU scheduling and resource management.
  • Experience with multi-tenant cloud environments.
  • Familiarity with Go-based cloud or platform services.
  • Experience with Infrastructure as Code and automated platform deployment.
  • Knowledge of distributed storage, networking or database troubleshooting.
  • Experience defining service-level indicators, objectives and operational reliability targets.
  • Experience working in sovereign, regulated or security-sensitive cloud environments.

Eligibility & Application Information

We welcome non-native Dutch speakers to apply. However, to be eligible, you must:

  • Have a valid work permit in the Netherlands. ( wo do offer sponsorship if needeed)
  • Reside in the Netherlands and be able to travel to the office in Leiden (near The Hague).
  • Be fluent in English. Dutch is not required.

Ready to make Europe’s sovereign AI cloud more reliable and operationally scalable?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Team Lead
Engineering Team Lead

Nebul • Leiden

On-site
EUR 90,000 - 120,000
Infrastructure Engineer
Infrastructure Engineer

Nebul • Leiden

On-site
EUR 55,000 - 75,000
DevOps Engineer – Go & Cloud Platform
DevOps Engineer – Go & Cloud Platform

Nebul • Leiden

On-site
EUR 75,000 - 110,000
SRE: Sovereign AI Cloud Platform Reliability
SRE: Sovereign AI Cloud Platform Reliability

Nebul • Leiden

On-site
EUR 70,000 - 110,000
Senior Full-Stack AI Engineer (LLMs & Agentic Systems)
Senior Full-Stack AI Engineer (LLMs & Agentic Systems)

Nebul • Leiden

On-site
EUR 90,000 - 130,000
Office near The Hague
Go Development Team Lead
Go Development Team Lead

Nebul • Leiden

On-site
EUR 70,000 - 90,000
Platform Engineer
Platform Engineer

Nebul • Leiden

On-site
EUR 60,000 - 100,000
Frontend Developer
Frontend Developer

Nebul • Leiden

On-site
EUR 60,000 - 80,000
Data Center - Service Delivery Manager
Data Center - Service Delivery Manager

Nebius • Amsterdam

On-site
EUR 75,000 - 95,000
Competitive salary
Comprehensive benefits package
Opportunities for professional growth
+1
Network Automation Engineer
Network Automation Engineer

Nebul • Leiden

On-site
EUR 70,000 - 110,000