Remote SRE: AI Infrastructure & GPU Cluster Reliability

Hamilton Barnes Associates Limited

Amstelveen

Remote

EUR 180,000 - 220,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

IPO Equity

Job summary

A seed-stage AI infrastructure company is seeking a Site Reliability Engineer for their European operations in a remote capacity. This role requires over 7 years of experience in SRE, DevOps, or Infrastructure Engineering, with expertise in Kubernetes and Slurm. Responsibilities include designing large-scale GPU clusters, building automation pipelines, and ensuring high availability of workloads. The position offers a competitive salary of €200,000+ gross annually, along with IPO equity benefits.

Qualifications

  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).

Responsibilities

  • Design, deploy, and maintain large-scale GPU clusters for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources.
  • Develop observability, alerting, and auto-healing systems for high-availability workloads.

Skills

Kubernetes
Slurm
Linux systems
Networking
GPU infrastructure
Python
Go
Bash

Tools

Prometheus
Grafana
Loki

Job description

A seed-stage AI infrastructure company is seeking a Site Reliability Engineer for their European operations in a remote capacity. This role requires over 7 years of experience in SRE, DevOps, or Infrastructure Engineering, with expertise in Kubernetes and Slurm. Responsibilities include designing large-scale GPU clusters, building automation pipelines, and ensuring high availability of workloads. The position offers a competitive salary of €200,000+ gross annually, along with IPO equity benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (EU Remote) - AI infrastructure
Site Reliability Engineer (EU Remote) - AI infrastructure

Hamilton Barnes Associates Limited • Amstelveen

Remote
EUR 180,000 - 220,000
IPO Equity
Remote SRE & Security Engineer (EU)
Remote SRE & Security Engineer (EU)

XLINQ • Randstad

Remote
EUR 60,000 - 80,000
Engineering Manager, SRE
Engineering Manager, SRE

Slashhash • Netherlands

Hybrid
EUR 65,000 - 147,000
Fully remote
Remote SRE Engineering Manager – Lead Reliability & Platform
Remote SRE Engineering Manager – Lead Reliability & Platform

Slashhash • Netherlands

Hybrid
EUR 65,000 - 147,000
Fully remote
Senior SRE: AI Platform & Cloud Reliability Lead
Senior SRE: AI Platform & Cloud Reliability Lead

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
Site Reliability Engineer
Site Reliability Engineer

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
SRE: AI Infrastructure (Early Career) — Cloud & Networking
SRE: AI Infrastructure (Early Career) — Cloud & Networking

Gewis • Amsterdam

Hybrid
EUR 20,000 - 31,000
Mentorship opportunities
Hands-on production experience
Growth potential
+1
Senior Linux SRE — Reliability Engineer, Scalable Systems
Senior Linux SRE — Reliability Engineer, Scalable Systems

The ICE Group • Amsterdam

On-site
EUR 70,000 - 90,000
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral • Amsterdam

On-site
EUR 90,000 - 130,000
Founding Applied AI SRE Engineer — Build & Secure
Founding Applied AI SRE Engineer — Build & Secure

Lindus Health • Amsterdam

On-site
EUR 90,000 - 130,000