Senior SRE: AI Inference Platform (GPU, Kubernetes)

Slashhash

Netherlands

Hybrid

EUR 90,000 - 130,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Slashhash is seeking a Senior Site Reliability Engineer in the Netherlands to own the reliability and observability of a large-scale AI inference platform. You will design telemetry pipelines, build monitoring for massive production signals, and optimize Kubernetes and GPU resource usage.

The role requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise with a focus on reliability, observability, and GPU-heavy workloads.

Qualifications

  • Senior level experience in site reliability or platform engineering with focus on AI workloads.
  • Strong understanding of Kubernetes-based architectures and reliability practices.
  • Proficient in observability pipelines and incident management processes.

Responsibilities

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and improve telemetry pipelines for metrics, logs, and traces.
  • Build monitoring capable of processing large production signals and extracting insights.
  • Configure and optimize Kubernetes for high availability and scalable GPU resources.
  • Tune autoscaling to improve GPU resource efficiency.
  • Develop Terraform modules embedding resilience into clusters and services.
  • Improve request-routing, retry, and failure-handling to minimize outages.
  • Develop automation to detect, isolate, and remediate incidents quickly.
  • Create and maintain runbooks for incident response and procedures.
  • Contribute to post-mortems and implement corrective actions to prevent recurrence.
  • Define reliability practices for high-throughput APIs with alerting and SLOs.
  • Investigate distributed-system failures across infra and app layers.
  • Optimize systems from kernel to application layer and support GPU workloads.
  • Collaborate with software engineers to embed reliability in product development.

Skills

Kubernetes
Prometheus
Grafana
Terraform
Python
Bash
Distributed Systems
SLOs
Infrastructure as Code
GPU Workloads
vLLM
Triton
Ray
MLOps
Incident Management

Tools

Terraform
Kubernetes
Prometheus

Job description

Slashhash is seeking a Senior Site Reliability Engineer in the Netherlands to own the reliability and observability of a large-scale AI inference platform. You will design telemetry pipelines, build monitoring for massive production signals, and optimize Kubernetes and GPU resource usage.

The role requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise with a focus on reliability, observability, and GPU-heavy workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Inference Platform Reliability
Senior SRE: AI Inference Platform Reliability

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Slashhash • Netherlands

Hybrid
EUR 90,000 - 130,000
Senior SRE: AI Platform & Cloud Reliability (Hybrid)
Senior SRE: AI Platform & Cloud Reliability (Hybrid)

Harnham • Rotterdam

Hybrid
EUR 70,000 - 110,000
Competitive salary
Hybrid working
Exposure to cloud & AI platforms
+1
Senior Compute Node SRE: AI Cloud Reliability Leader
Senior Compute Node SRE: AI Cloud Reliability Leader

Jobgether • Netherlands

On-site
EUR 110,000 - 170,000
Competitive compensation
Career growth and continuous learning
Ownership and flexibility in work
+4
Senior SRE: AI Platform & Cloud Reliability Lead
Senior SRE: AI Platform & Cloud Reliability Lead

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harnham • Rotterdam

Hybrid
EUR 70,000 - 110,000
Competitive salary
Hybrid working
Exposure to cloud & AI platforms
+1
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1
Senior SRE - Hardware Automation for Scalable AI Infra
Senior SRE - Hardware Automation for Scalable AI Infra

Coinscapture • Netherlands

Remote
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Infra Agent Systems Engineer: AI Agents & Platform
Senior Infra Agent Systems Engineer: AI Agents & Platform

Together • Amsterdam

Hybrid
EUR 90,000 - 130,000
Lead AI Infra & SRE Engineering Team
Lead AI Infra & SRE Engineering Team

Together AI • Amsterdam

On-site
EUR 110,000 - 150,000
Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
+10