Site Reliability Engineer: AI Cloud Platform & Automation

Nebul

Leiden

On-site

EUR 90,000 - 120,000

Full time

24 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Nebul is building Europe’s sovereign AI cloud and is seeking a Site Reliability Engineer to keep the platform stable, observable and scalable. You’ll own incidents, tackle complex problems and reduce reliance on senior engineers by automation.

Working with Kubernetes, NVIDIA GPU infrastructure and services written in Go and Python, you’ll drive reliability, observability and proactive improvements across production systems within a collaborative team in Leiden.

Qualifications

  • Proven experience as a Site Reliability/DevOps/Platform engineer.
  • Hands-on production exposure to Kubernetes.
  • Strong Linux and infrastructure troubleshooting skills.
  • Experience with Go or Python services.
  • Ability to lead complex troubleshooting sessions.
  • Experience with monitoring, logging, metrics, and alerting.

Responsibilities

  • Maintain Nebul's AI cloud platform for reliability and scalability.
  • Own incidents and coordinate resolution with senior engineers.
  • Investigate issues across Kubernetes, GPU infrastructure, and services in Go/Python.
  • Automate repetitive operational tasks.
  • Improve platform reliability with runbooks and documentation.
  • Support deployments and platform changes.
  • Enhance observability and failure handling.

Skills

Kubernetes
Go
Python
Linux
NVIDIA GPU
Monitoring
Observability
Incident response
Automation
Cloud-native
Networking

Tools

Go
Python

Job description

Nebul is building Europe’s sovereign AI cloud and is seeking a Site Reliability Engineer to keep the platform stable, observable and scalable. You’ll own incidents, tackle complex problems and reduce reliance on senior engineers by automation.

Working with Kubernetes, NVIDIA GPU infrastructure and services written in Go and Python, you’ll drive reliability, observability and proactive improvements across production systems within a collaborative team in Leiden.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – AI Cloud Platform
Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden

On-site
EUR 90,000 - 120,000
SRE: AI Infrastructure (Early Career) — Cloud & Networking
SRE: AI Infrastructure (Early Career) — Cloud & Networking

Gewis • Amsterdam

Hybrid
EUR 20,000 - 31,000
Mentorship opportunities
Hands-on production experience
Growth potential
+1
Senior Full-Stack AI Engineer - Production AI Apps
Senior Full-Stack AI Engineer - Production AI Apps

Nebul • Leiden

On-site
EUR 90,000 - 130,000
Office near The Hague
Backend Engineer (GO)
Backend Engineer (GO)

Nebul • Leiden

On-site
EUR 90,000 - 130,000
Go Backend Engineer — Cloud Platform & Kubernetes
Go Backend Engineer — Cloud Platform & Kubernetes

Nebul • Leiden

On-site
EUR 90,000 - 130,000
Site Reliability Engineer (SRE) AI Infrastructure (Early Career) at Nebius
Site Reliability Engineer (SRE) AI Infrastructure (Early Career) at Nebius

Gewis • Amsterdam

Hybrid
EUR 20,000 - 31,000
Mentorship opportunities
Hands-on production experience
Growth potential
+1
Senior Software Engineer - Serverless AI Platform
Senior Software Engineer - Serverless AI Platform

Nebius B.V. • Amsterdam

Hybrid
EUR 110,000 - 140,000
Hybrid work arrangement
Infrastructure Engineer
Infrastructure Engineer

Nebul • Leiden

On-site
EUR 55,000 - 75,000
Golang Backend Software Engineer
Golang Backend Software Engineer

Nebul • Leiden

On-site
EUR 70,000 - 110,000
Sovereign Cloud Network Automation Engineer
Sovereign Cloud Network Automation Engineer

Nebul • Leiden

On-site
EUR 70,000 - 110,000