Senior Compute Node SRE: AI Cloud Reliability Leader

Jobgether

Netherlands

On-site

EUR 110,000 - 170,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Career growth and continuous learning
Ownership and flexibility in work
Collaborative engineering environment
Impactful AI and cloud infrastructure
Exposure to large-scale compute and OS
International engineering teams

Job summary

Jobgether, on behalf of a partner, is seeking a Senior Site Reliability Engineer based in the Netherlands to own compute node reliability across a large-scale cloud platform. The role emphasizes Linux, virtualization, and containerized workloads.

You will help shape reliability practices through strong monitoring, incident response, root-cause analysis, and postmortems while collaborating with platform, kernel, hypervisor, GPU, and infrastructure teams to improve system design and operability.

Qualifications

  • Significant experience in Site Reliability Engineering and Linux infra.
  • Deep knowledge of Linux user and kernel space.
  • Hands-on virtualization expertise with QEMU/KVM.
  • Experience with containers, namespaces, and cgroups.
  • Strong incident response, postmortems, and observability focus.

Responsibilities

  • Ensure compute node reliability and performance for running VMs.
  • Debug Linux systems across user and kernel space.
  • Analyze dependencies and trade-offs across OS and infra layers.
  • Troubleshoot CPU, memory, NUMA, cgroups, and scheduling issues.
  • Work with virtualization technologies (QEMU/KVM) and Linux-native tech.
  • Design and evolve observability with metrics, logs, traces, alerts, SLIs, and SLOs.
  • Lead or contribute to incident response and postmortems.
  • Collaborate with platform, kernel/hypervisor, GPU, and infrastructure teams.

Skills

Linux expertise
SRE discipline
Kubernetes internals
QEMU/KVM
Containers
Observability
Incident response
Root-cause analysis
Perf debugging
Kernel debugging tools

Tools

Perf
eBPF
ftrace
strace

Job description

Jobgether, on behalf of a partner, is seeking a Senior Site Reliability Engineer based in the Netherlands to own compute node reliability across a large-scale cloud platform. The role emphasizes Linux, virtualization, and containerized workloads.

You will help shape reliability practices through strong monitoring, incident response, root-cause analysis, and postmortems while collaborating with platform, kernel, hypervisor, GPU, and infrastructure teams to improve system design and operability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Netherlands

On-site
EUR 110,000 - 170,000
Competitive compensation
Career growth and continuous learning
Ownership and flexibility in work
+4
Senior SRE: AI Platform & Cloud Reliability (Hybrid)
Senior SRE: AI Platform & Cloud Reliability (Hybrid)

Harnham • Rotterdam

Hybrid
EUR 70,000 - 110,000
Competitive salary
Hybrid working
Exposure to cloud & AI platforms
+1
Senior SRE: AI Platform & Cloud Reliability Lead
Senior SRE: AI Platform & Cloud Reliability Lead

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harnham • Rotterdam

Hybrid
EUR 70,000 - 110,000
Competitive salary
Hybrid working
Exposure to cloud & AI platforms
+1
Senior SRE: AI Inference Platform Reliability
Senior SRE: AI Inference Platform Reliability

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1
Senior SRE: AI Inference Platform (GPU, Kubernetes)
Senior SRE: AI Inference Platform (GPU, Kubernetes)

Slashhash • Netherlands

Hybrid
EUR 90,000 - 130,000
Senior SRE - Hybrid Amsterdam - Equity
Senior SRE - Hybrid Amsterdam - Equity

DataSnipper • Amsterdam

Hybrid
EUR 90,000 - 140,000
Equity
Pension
Vacation days
+7
Lead AI Infra & SRE Engineering Team (Amsterdam)
Lead AI Infra & SRE Engineering Team (Amsterdam)

Together AI • Amsterdam

On-site
EUR 110,000 - 160,000
Senior SRE - Hardware Automation for Scalable AI Infra
Senior SRE - Hardware Automation for Scalable AI Infra

Coinscapture • Netherlands

Remote
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Lead AI Infra & SRE Engineering Team
Lead AI Infra & SRE Engineering Team

Together AI • Amsterdam

On-site
EUR 110,000 - 150,000
Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
+10