Senior SRE - AI GPU Infra & HPC Orchestrator (Remote EU)

Hamilton Barnes ?

España

Presencial

EUR 120.000 - 190.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Hamilton Barnes is seeking a Senior Site Reliability Engineer to own reliability, performance, and automation for a hyperscale GPU-centric infrastructure. You will shape operations for thousands of nodes and work across Slurm and Kubernetes environments.

Join a stealth-mode data center startup focused on AI workloads, with international remote scope and opportunities to influence high-availability GPU clusters at scale.

Formación

  • 7+ years of SRE/DevOps experience supporting large-scale compute environments.
  • Strong hands-on Kubernetes and Slurm for cluster orchestration.
  • Deep Linux, networking and GPU infrastructure knowledge.
  • Automation in Python/Go/Bash for tooling and performance tuning.
  • Experience with observability stacks and incident response frameworks.
  • Familiarity with HPC/AI training infrastructure at scale.
  • Background in reliability engineering or distributed systems is a plus.

Responsabilidades

  • Design, deploy, and maintain large-scale GPU clusters for training and inference.
  • Build automation for provisioning, scaling, and monitoring compute resources.
  • Develop observability, alerting, and auto-healing systems.
  • Collaborate with ML, networking, and platform teams to optimise scheduling.
  • Implement infrastructure-as-code, CI/CD, and reliability standards.
  • Diagnose bottlenecks and drive improvements in reliability and latency.

Conocimientos

SRE/DevOps experience
Kubernetes
Slurm
Linux
Python/Go/Bash
Observability
HPC/AI infra

Herramientas

Prometheus
Grafana
Loki
CI/CD

Descripción del empleo

Hamilton Barnes is seeking a Senior Site Reliability Engineer to own reliability, performance, and automation for a hyperscale GPU-centric infrastructure. You will shape operations for thousands of nodes and work across Slurm and Kubernetes environments.

Join a stealth-mode data center startup focused on AI workloads, with international remote scope and opportunities to influence high-availability GPU clusters at scale.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior AI Infra & Platform SRE — Remote EU
Senior AI Infra & Platform SRE — Remote EU

Front Door Defense • Barcelona

A distancia
EUR 90.000 - 130.000
Remote work flexibility
Competitive compensation
Career growth opportunities
Senior Network Engineer — AI GPU Cloud & Low-Latency DCs
Senior Network Engineer — AI GPU Cloud & Low-Latency DCs

Hamilton Barnes ? • España

Presencial
EUR 80.000 - 110.000
Senior SRE - AI Cloud Infra, Kubernetes & GPUs
Senior SRE - AI Cloud Infra, Kubernetes & GPUs

Mirantis • Bellprat

Presencial
EUR 90.000 - 120.000
Competitive compensation package
Professional development and training
Conference attendance and tech talks
+1
Remote-First Senior SRE: Scalable AI Infra
Remote-First Senior SRE: Scalable AI Infra

Runware • España

Presencial
EUR 70.000 - 100.000
Generous paid time off
Stock options
Remote-first setup
+3
Senior SRE: Cloud Reliability & Automation (Remote)
Senior SRE: Cloud Reliability & Automation (Remote)

SciSure • España

Presencial
EUR 90.000 - 125.000
Fully remote position
30 days vacation per year
Support for courses and conferences
+1
Senior GPU Cloud Infra & Deployment Engineer
Senior GPU Cloud Infra & Deployment Engineer

Jobgether • España

Presencial
EUR 90.000 - 150.000
Remote-friendly Europe-based remote
GPU cloud exposure
NVIDIA GPUs
Senior SRE - Remote, Build Reliable, Scalable Platform
Senior SRE - Remote, Build Reliable, Scalable Platform

Randstad (Schweiz) AG • Madrid

Presencial
EUR 70.000 - 100.000
Health Insurance
Paid Time Off (PTO)
Paid Holidays
+2
Senior SRE: AI Platform & Cloud Reliability
Senior SRE: AI Platform & Cloud Reliability

Doist • Barcelona

Híbrido
EUR 85.000 - 120.000
Hybrid onboarding
Health insurance
Professional development budget
+4
Senior SRE — Global Remote, Scalable Systems
Senior SRE — Global Remote, Scalable Systems

Fountain • Madrid

Presencial
EUR 70.000 - 110.000
Competitive health plans
Retirement plan
Flexible vacation policy
+2
Head of SRE - AI Reliability & Infra Leader (Remote Europe)
Head of SRE - AI Reliability & Infra Leader (Remote Europe)

Jobgether • España

Presencial
EUR 120.000 - 180.000
Fully remote within European time zone
European time zone alignment
Significant ownership in engineering &