Site Reliability Engineer (EU Remote) - AI infrastructure

Hamilton Barnes Associates Limited

Amstelveen

Remote

EUR 180,000 - 220,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

IPO Equity

Job summary

A seed-stage AI infrastructure company is seeking a Site Reliability Engineer for their European operations in a remote capacity. This role requires over 7 years of experience in SRE, DevOps, or Infrastructure Engineering, with expertise in Kubernetes and Slurm. Responsibilities include designing large-scale GPU clusters, building automation pipelines, and ensuring high availability of workloads. The position offers a competitive salary of €200,000+ gross annually, along with IPO equity benefits.

Qualifications

  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).

Responsibilities

  • Design, deploy, and maintain large-scale GPU clusters for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources.
  • Develop observability, alerting, and auto-healing systems for high-availability workloads.

Skills

Kubernetes
Slurm
Linux systems
Networking
GPU infrastructure
Python
Go
Bash

Tools

Prometheus
Grafana
Loki

Job description

Overview

Join a seed-stage AI infrastructure company building large-scale training and inference platforms previously accessible only to hyperscalers. The business began with a single managed GPU cluster that reached capacity almost immediately and has since expanded into a global platform spanning infrastructure, networking, and orchestration.

They are now seeking a Site Reliability Engineer to join their Eu operations in this remote role. This is ideal if you have 7 years of experience in SRE, DevOps, or Infrastructure Engineering roles and have had exposure to supporting large-scale compute environments.

If you are interested in this exciting opportunity, get in touch and apply today!

Responsibilities
  • Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
  • Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
  • Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilisation, and data flow.
  • Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
  • Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.
Qualifications
  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
  • Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
  • Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
  • Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
  • Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.
Benefits
  • IPO Equity
Salary
  • €200,000+ gross per year
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE: AI Infrastructure & GPU Cluster Reliability
Remote SRE: AI Infrastructure & GPU Cluster Reliability

Hamilton Barnes Associates Limited • Amstelveen

Remote
EUR 180,000 - 220,000
IPO Equity
Site Selection Manager - AI Infrastructure
Site Selection Manager - AI Infrastructure

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 117,000 - 143,000
Competitive salary
Bonus
Comprehensive benefits package
+1
GPU Cluster Architect - Data Center
GPU Cluster Architect - Data Center

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 180,000 - 220,000
Bonus scheme
Company shares
Flexible remote working
Senior Sales Consultant (EU wide) - Hosting
Senior Sales Consultant (EU wide) - Hosting

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 72,000 - 120,000
Remote work with international travel
Equity opportunities
Competitive salary + bonus
+1
Site Reliability Engineer
Site Reliability Engineer

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
Project Development Manager - AI Infrastructure
Project Development Manager - AI Infrastructure

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 117,000 - 143,000
Competitive salary
Bonus potential
Progression opportunities
Regional Data Center Operations Manager - Systems Integrator
Regional Data Center Operations Manager - Systems Integrator

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 126,000 - 154,000
Clear progression and leadership development opportunities
Flexible working arrangements
Bonus scheme
Senior Hardware Engineer - Systems Integrator
Senior Hardware Engineer - Systems Integrator

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 108,000 - 132,000
Competitive salary and benefits package
Opportunity to work on cutting edge GPU and AI infrastructure
Collaborative engineering driven culture
HPC Cluster Engineer - AI Infrastructure
HPC Cluster Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Amsterdam

On-site
EUR 70,000 - 90,000
Competitive salary and full benefits package
Opportunities for professional growth
Hybrid work environment
+1
AI Benchmark Engineer
AI Benchmark Engineer

Hamilton Barnes ? • Netherlands

Hybrid
EUR 85,000 - 145,000