Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.
Technical (Must-have)
- Kubernetes
- Prometheus
- Grafana
- Terraform
- Python
- Bash
- Distributed Systems
- SLOs
- Infrastructure as Code
- GPU Workloads
- vLLM
- Triton
- Ray
- MLOps
- Incident Management
Soft Skills
- Troubleshooting
- Collaboration
- Proactive Mindset
- Ownership
- Continuous Improvement
- Root Cause Analysis
Technical (Nice-to-have)
- Model Hosting
- AI Infrastructure
- Machine Learning Platforms
Key Responsibilities
- Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
- Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
- Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
- Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
- Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
- Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
- Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
- Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
- Create, maintain, and improve runbooks for incident response and operational procedures.
- Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
- Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
- Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
- Investigate distributed-system failures and performance issues across infrastructure and application layers.
- Optimize systems from the kernel and infrastructure layer through to the application layer.
- Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
- Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
- Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
- Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
Kubernetes, Prometheus, Grafana, Terraform, Python, Bash, Distributed Systems, SLOs, Infrastructure as Code, GPU Workloads, vLLM, Triton, Ray, MLOps, Incident Management