Senior Site Reliability Engineer (SRE) – Travel – Amsterdam
Duration: 6-Month Contract
Start: ASAP
Hybrid: 2 days a week (on-call 1/7)
My client is looking for an experienced Site Reliability Engineer to join a high-performing platform team responsible for operating and evolving a mission-critical event processing platform that handles billions of events every day.
This is an exciting opportunity to work on large-scale distributed systems, drive operational excellence, and support a major cloud migration initiative while ensuring the reliability, scalability, and performance of a business-critical platform.
What You'll Be Doing
- Own end-to-end reliability of production services.
- Lead incident response, root cause analysis, post-mortems, and remediation activities.
- Improve observability through monitoring, alerting, logging, tracing, and dashboards.
- Build and enhance CI/CD pipelines and Infrastructure-as-Code solutions.
- Drive automation initiatives to reduce operational toil and improve platform efficiency.
- Support performance testing, capacity planning, and scalability initiatives.
- Contribute to the migration of a large-scale event streaming platform to a cloud-native architecture.
- Collaborate with software engineers and platform teams to improve reliability, resilience, and operational maturity.
- Participate in a shared on-call rotation and help maintain high service availability.
What We're Looking For
- Proven experience as a Site Reliability Engineer, SRE, Platform Engineer, or DevOps Engineer in high-scale production environments.
- Strong hands-on experience with:
- Java
- Kafka
- AWS
- Terraform, Helm, GitOps, or similar Infrastructure-as-Code tools
- CI/CD pipelines and deployment automation
- Experience with modern observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar.
- Strong background in incident management, production troubleshooting, and reliability engineering.
- Experience operating and scaling distributed systems handling high transaction or event volumes.
- Excellent communication skills and the ability to work effectively across engineering teams.
Nice to Have
- Experience with Confluent Cloud.
- Experience migrating Kafka workloads from on-premises environments to cloud platforms.
- Knowledge of large-scale event-driven architectures and stream processing systems.
- Experience optimising platform performance, scalability, and operational costs.