Senior Site Reliability Engineer (SRE & AI Platform Operations)
Rotterdam, 5 days in office
Up to €95k annually + benefits
This is an opportunity to take ownership of reliability, observability, cloud infrastructure, and AI platform operations within a fast-moving technology environment. You will play a key role in building the guardrails, automation, and operational excellence that enable engineering teams to innovate at pace while maintaining stability and performance.
The Company
They are a technology-driven organisation undergoing significant platform transformation, modernising core systems and investing heavily in cloud-native architecture and AI-enabled capabilities. Their engineering culture is focused on innovation, automation, and continuous improvement. You will join a collaborative environment where technical expertise is valued and where your work will have a direct impact on business performance and customer experience.
The Role
- Define and manage Service Level Objectives (SLOs), SLIs, and error budget policies across critical services.
- Lead initiatives across observability, distributed tracing, monitoring, and incident response.
- Improve deployment safety through CI/CD best practices, automated rollbacks, and progressive delivery techniques.
- Own AI platform operations, including runtime performance, scalability, reliability, and cost optimisation.
- Drive FinOps initiatives across cloud infrastructure, identifying opportunities to improve efficiency and manage costs.
- Lead incident management activities and develop automated safeguards to prevent recurring issues.
- Enhance resilience through capacity planning, load testing, disaster recovery planning, and service reliability improvements.
- Manage infrastructure through Infrastructure as Code and cloud automation practices.
- Build tooling, runbooks, and self-service capabilities that improve the developer experience and reduce operational overhead.
Your Skills & Experience
- Strong commercial experience operating high-traffic, distributed production systems.
- Deep knowledge of Site Reliability Engineering principles, monitoring, and platform operations.
- Hands-on experience with Google Cloud Platform and Infrastructure as Code using Terraform.
- Strong troubleshooting capabilities across Linux environments, cloud infrastructure, containers, databases, and distributed systems.
- Experience with observability tooling, monitoring platforms, and distributed tracing.
- Familiarity with modern software engineering environments and cloud-native architectures.
- Understanding of AI workloads, LLM integrations, API-driven services, or automated processing pipelines.
- A proactive, solutions-focused approach with a passion for automation and operational excellence.
What They Offer
- Competitive salary and benefits package.
- The opportunity to shape platform reliability and AI operations at scale.
- Real ownership and influence across engineering and infrastructure strategy.
- An international and collaborative working environment.
- Clear opportunities for professional growth and career progression.
- The chance to work on cutting-edge cloud and AI technologies.