Senior Site Reliability Engineer — Token Factory (Inference Platform)

Slashhash

Netherlands

Hybrid

EUR 90,000 - 130,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Slashhash is seeking a Senior Site Reliability Engineer in the Netherlands to own the reliability and observability of a large-scale AI inference platform. You will design telemetry pipelines, build monitoring for massive production signals, and optimize Kubernetes and GPU resource usage.

The role requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise with a focus on reliability, observability, and GPU-heavy workloads.

Qualifications

  • Senior level experience in site reliability or platform engineering with focus on AI workloads.
  • Strong understanding of Kubernetes-based architectures and reliability practices.
  • Proficient in observability pipelines and incident management processes.

Responsibilities

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and improve telemetry pipelines for metrics, logs, and traces.
  • Build monitoring capable of processing large production signals and extracting insights.
  • Configure and optimize Kubernetes for high availability and scalable GPU resources.
  • Tune autoscaling to improve GPU resource efficiency.
  • Develop Terraform modules embedding resilience into clusters and services.
  • Improve request-routing, retry, and failure-handling to minimize outages.
  • Develop automation to detect, isolate, and remediate incidents quickly.
  • Create and maintain runbooks for incident response and procedures.
  • Contribute to post-mortems and implement corrective actions to prevent recurrence.
  • Define reliability practices for high-throughput APIs with alerting and SLOs.
  • Investigate distributed-system failures across infra and app layers.
  • Optimize systems from kernel to application layer and support GPU workloads.
  • Collaborate with software engineers to embed reliability in product development.

Skills

Kubernetes
Prometheus
Grafana
Terraform
Python
Bash
Distributed Systems
SLOs
Infrastructure as Code
GPU Workloads
vLLM
Triton
Ray
MLOps
Incident Management

Tools

Terraform
Kubernetes
Prometheus

Job description

Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.

Technical (Must-have)
  • Kubernetes
  • Prometheus
  • Grafana
  • Terraform
  • Python
  • Bash
  • Distributed Systems
  • SLOs
  • Infrastructure as Code
  • GPU Workloads
  • vLLM
  • Triton
  • Ray
  • MLOps
  • Incident Management
Soft Skills
  • Troubleshooting
  • Collaboration
  • Proactive Mindset
  • Ownership
  • Continuous Improvement
  • Root Cause Analysis
Technical (Nice-to-have)
  • Model Hosting
  • AI Infrastructure
  • Machine Learning Platforms
Key Responsibilities
  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
  • Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
  • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
  • Create, maintain, and improve runbooks for incident response and operational procedures.
  • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
  • Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
  • Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
  • Investigate distributed-system failures and performance issues across infrastructure and application layers.
  • Optimize systems from the kernel and infrastructure layer through to the application layer.
  • Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
  • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
  • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
  • Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.

Kubernetes, Prometheus, Grafana, Terraform, Python, Bash, Distributed Systems, SLOs, Infrastructure as Code, GPU Workloads, vLLM, Triton, Ray, MLOps, Incident Management

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1
Senior SRE: AI Inference Platform Reliability
Senior SRE: AI Inference Platform Reliability

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harnham • Rotterdam

Hybrid
EUR 70,000 - 110,000
Competitive salary
Hybrid working
Exposure to cloud & AI platforms
+1
Site Reliability Engineer – AI Cloud Platform
Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden

On-site
EUR 90,000 - 120,000
Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Netherlands

On-site
EUR 110,000 - 170,000
Competitive compensation
Career growth and continuous learning
Ownership and flexibility in work
+4
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

United States Digital Space LLC • Amsterdam

Hybrid
EUR 100,000 - 180,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

Slashhash • Amsterdam

Hybrid
EUR 120,000 - 150,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together • Amsterdam

Hybrid
EUR 90,000 - 130,000
Site Reliability Engineer: AI Cloud Platform & Automation
Site Reliability Engineer: AI Cloud Platform & Automation

Nebul • Leiden

On-site
EUR 90,000 - 120,000
Senior AI Infrastructure Engineer Amsterdam
Senior AI Infrastructure Engineer Amsterdam

Together Computer Inc • Amsterdam

Hybrid
EUR 120,000 - 150,000