Salary: £68,000 - 108,000 per year
Requirements:
- 3+ years experience in SRE, Platform, or DevOps roles within production environments.
- Strong Kubernetes operational experience (on-prem and AWS EKS).
- Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows.
- Deep understanding of observability and telemetry (monitoring, logging, tracing).
- Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD.
- Scripting proficiency in Python, Bash, or Go.
- Proven ability to balance cost efficiency with reliability and performance.
- Excellent communication skills and the ability to work effectively across multiple teams.
- Experience running chaos engineering experiments is desirable.
- Exposure to high-throughput, low-latency systems is desirable.
- FinOps knowledge or cost management practices is desirable.
- AWS certifications (e.g., Solutions Architect, DevOps Engineer) are desirable.
Responsibilities:
- Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions.
- Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively.
- Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling, and improve infrastructure efficiency.
- Improve production readiness via pre-deployment checks, post-release validation, and robust platform guardrails.
- Introduce and run chaos engineering experiments to strengthen resilience and recovery.
- Automate operational processes to reduce manual intervention and toil across the stack.
- Support major incident response, root-cause analysis, and continual improvement actions.
- Collaborate cross-functionally to raise standards for stability, security, performance, and compliance.
Technologies:
- AWS
- Architect
- Bash
- CI/CD
- DevOps
- GitOps
- Support
- Kubernetes
- Python
- Security
- Terraform
- Cloud
- Incident Management
More:
We are a premier provider of high-volume software solutions for the global iGaming and predictive analytics sector, with a footprint spanning the USA, UK, and Europe. We partner with industry leaders to engineer sophisticated platforms for sports wagering, prize-based systems, and complex market simulation environments. Our vision is to lead the evolution of interactive technology through intelligent, data-driven architecture that ensures seamless user experiences. We operate with a culture of teamwork, transparency, and technical excellence. This is a full-time permanent Site Reliability Engineer role based in London City on a hybrid basis, with 3 days per week on-site, and a salary of 90,000 per annum.
last updated 36 week of 2026