The Team
OpenBetis a global leader in betting and gaming entertainment, trusted by over 200 partners to create memorable winning moments for millions of players worldwide. From processing bets during iconic events like theFIFA World CupandSuper Bowlto pioneering next-gen products likeBetBuilder, we continuously redefine the player experience with high-quality content,cutting-edgetechnology, and advanced player protection tools.
For over 25 years, our unbeatable platform has powered the most recognizable betting brands, ensuring peak performance with100% uptime, unmatched scale, and speed. With 85 licenses, 20 World Lottery Association operators on our customer roster, and a team of 1,200+ experts across 14 countries, weremainat the heart of the industry.
The Goal
What will your role be?
We're looking for an SRE Manager to lead our Site Reliability Engineering team, with a particular focus on observability and performance engineering. You'll build and lead the team responsible for how we see, measure, and understand the health of our platform — setting the standards for monitoring, alerting, and performance testing that keep our systems running reliably. You'll balance hands-on technical leadership with people management, working closely with DevOps, Infrastructure, and Database Engineering to raise the bar on reliability across the estate.
What you’ll be doing
- Lead and grow the SRE team, providing day-to-day people management, coaching, and career development for a group of reliability and performance engineers
- Own the observability strategy across the platform — metrics, logging, tracing, and alerting — equipping DevOps, Infrastructure, and Engineering teams with self-service dashboards and insight rather than gatekeeping the data
- Own the evaluation, budget, and vendor relationship for observability tooling, balancing capability, cost, and operational fit
- Explore and adopt AI-assisted observability capabilities — anomaly detection, predictive alerting, and AI-driven root-cause analysis — to help the team spot issues earlier and resolve them faster
- Own the performance engineering practice, including load testing, capacity planning, and performance benchmarking
- Define and drive SLIs, SLOs, and error budgets for critical platform services, working with Engineering to embed reliability targets into delivery
- Act as an escalation point for major incidents, contributing root cause analysis and ensuring learnings feed back into monitoring, alerting, and performance improvements
- Partner with DevOps, Infrastructure, and Database Engineering to close observability gaps across the hybrid on-premises and cloud estate
- Ensure the team maintains clear runbooks and playbooks for common failure scenarios, making reliability knowledge repeatable rather than dependent on individual expertise
- Champion a proactive reliability culture, shifting the team from reactive firefighting toward prevention through better tooling, automation, and standards
The Player
- Proven experience leading a Site Reliability Engineering or Performance Engineering team, including direct people management responsibility
- Deep hands-on background in observability — metrics, logging, tracing, and alerting
- Strong experience with performance engineering practices, including load testing, capacity planning, and performance benchmarking
- A track record of defining and operationalising SLIs, SLOs, and error budgets in a production environment
- Experience acting as an escalation point for major incident response, contributing to root cause analysis through to concrete reliability improvements
- Comfort working across hybrid on-premises and cloud environments, ideally within a high-availability, transaction-heavy domain
- Hands-on AWS experience, with a focus on cost optimization and leveraging AI-powered tools to drive efficiency across cloud infrastructure
- Working knowledge of Kubernetes and Rancher, enough to operate confidently across our hybrid on-premises and cloud platforms
- An AI-native approach to observability, with exposure to AI or LLM-assisted tools for anomaly detection and root-cause analysis a plus
- Strong communication skills, with the ability to translate technical reliability data into insight for non-technical stakeholders
- A coaching mindset, with genuine enthusiasm for developing engineers and building a strong team culture
What’sthe Score?
WhyOpenBet?
- The Playground:Join a team of innovators, disruptors, andgame-changerswho are reshaping the future of betting and gaming.
- The Mission:Be part of a mission-driven organizationthat'scommitted to revolutionizing the way the world plays.
- The Impact:Make a real impact on the world stage, leavinga lasting legacythat transcends boundaries and inspires generations to come.
- The Culture:Immerse yourself in a culture of creativity, collaboration, and curiosity, where every idea is welcomed, every voice is heard, and every dream is encouraged.
- The Future:Join us on the journey to build the future of betting and gaming, one game-changing innovation at a time.
What we can offer YOU:
- Attractive benefits, an open and supportive environment as well as a modern and exciting workplace
- The opportunity to interact with global teams on a regular basis as you and our business continues to develop & grow
- Tangible and genuine development - atOpenBet, you can take your career where you want it to go!
- And ifthat’snot enough;enjoyflexibleworkingwhilst we provide you with theguidanceanddevelopmentskillsyou need to progress andenhance your career
- We have a collaborative office environment with our team members in office 2 days per week.
AtOpenBet, we celebrate diversity and believe in creating an inclusive environment where every voice is valued and respected.We'recommitted to building a team that reflects the rich tapestry of humanity, embracing individuals from allwalks of life, backgrounds, and identities. Join us in shaping the future of iGaming, where diversityisn'tjust celebrated— it'scelebrated.