SRE Manager

OpenBet

Kentucky

Hybrid

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive benefits
Global collaboration
Flexible working

Job summary

OpenBet, a global leader in betting and gaming, is seeking a seasoned SRE Manager to lead our Site Reliability Engineering team with a focus on observability and performance engineering. You will set monitoring standards, coach engineers, and partner across DevOps, Infrastructure, and Database Engineering to raise reliability across the hybrid estate.

You will guide the observability strategy, own SLI/SLO definitions, and drive AI-assisted approaches to incident detection and root cause

Qualifications

  • Proven experience leading an SRE or Performance Engineering team.
  • Strong observability expertise: metrics, logging, tracing, alerting.
  • Hands-on AWS experience with cost optimization and AI-powered tooling beneficial.
  • Experience defining and operationalising SLIs, SLOs, and error budgets.

Responsibilities

  • Lead and grow the SRE team with people management, coaching, and career development.
  • Own observability strategy across the platform: metrics, logging, tracing, alerting, dashboards.
  • Manage evaluation, budget, and vendor relationships for observability tooling.
  • Explore AI-assisted observability capabilities for anomaly detection and predictive alerting.
  • Own performance engineering: load testing, capacity planning, benchmarking.
  • Define and drive SLIs, SLOs, and error budgets with engineering teams.
  • Act as escalation point for major incidents and perform root cause analysis.
  • Collaborate with DevOps, Infrastructure, and Database Engineering to close gaps across hybrid estate.
  • Maintain runbooks and playbooks for reliable incident response and prevention.

Skills

SRE leadership
Observability
Performance engineering
AWS
Kubernetes
Rancher
AI-powered observability
Verbal/written communication
Incident response

Tools

Kubernetes
Rancher

Job description

The Team
OpenBetis a global leader in betting and gaming entertainment, trusted by over 200 partners to create memorable winning moments for millions of players worldwide. From processing bets during iconic events like theFIFA World CupandSuper Bowlto pioneering next-gen products likeBetBuilder, we continuously redefine the player experience with high-quality content,cutting-edgetechnology, and advanced player protection tools.

For over 25 years, our unbeatable platform has powered the most recognizable betting brands, ensuring peak performance with100% uptime, unmatched scale, and speed. With 85 licenses, 20 World Lottery Association operators on our customer roster, and a team of 1,200+ experts across 14 countries, weremainat the heart of the industry.

The Goal
What will your role be?

We're looking for an SRE Manager to lead our Site Reliability Engineering team, with a particular focus on observability and performance engineering. You'll build and lead the team responsible for how we see, measure, and understand the health of our platform — setting the standards for monitoring, alerting, and performance testing that keep our systems running reliably. You'll balance hands-on technical leadership with people management, working closely with DevOps, Infrastructure, and Database Engineering to raise the bar on reliability across the estate.

What you’ll be doing
  • Lead and grow the SRE team, providing day-to-day people management, coaching, and career development for a group of reliability and performance engineers
  • Own the observability strategy across the platform — metrics, logging, tracing, and alerting — equipping DevOps, Infrastructure, and Engineering teams with self-service dashboards and insight rather than gatekeeping the data
  • Own the evaluation, budget, and vendor relationship for observability tooling, balancing capability, cost, and operational fit
  • Explore and adopt AI-assisted observability capabilities — anomaly detection, predictive alerting, and AI-driven root-cause analysis — to help the team spot issues earlier and resolve them faster
  • Own the performance engineering practice, including load testing, capacity planning, and performance benchmarking
  • Define and drive SLIs, SLOs, and error budgets for critical platform services, working with Engineering to embed reliability targets into delivery
  • Act as an escalation point for major incidents, contributing root cause analysis and ensuring learnings feed back into monitoring, alerting, and performance improvements
  • Partner with DevOps, Infrastructure, and Database Engineering to close observability gaps across the hybrid on-premises and cloud estate
  • Ensure the team maintains clear runbooks and playbooks for common failure scenarios, making reliability knowledge repeatable rather than dependent on individual expertise
  • Champion a proactive reliability culture, shifting the team from reactive firefighting toward prevention through better tooling, automation, and standards
The Player
  • Proven experience leading a Site Reliability Engineering or Performance Engineering team, including direct people management responsibility
  • Deep hands-on background in observability — metrics, logging, tracing, and alerting
  • Strong experience with performance engineering practices, including load testing, capacity planning, and performance benchmarking
  • A track record of defining and operationalising SLIs, SLOs, and error budgets in a production environment
  • Experience acting as an escalation point for major incident response, contributing to root cause analysis through to concrete reliability improvements
  • Comfort working across hybrid on-premises and cloud environments, ideally within a high-availability, transaction-heavy domain
  • Hands-on AWS experience, with a focus on cost optimization and leveraging AI-powered tools to drive efficiency across cloud infrastructure
  • Working knowledge of Kubernetes and Rancher, enough to operate confidently across our hybrid on-premises and cloud platforms
  • An AI-native approach to observability, with exposure to AI or LLM-assisted tools for anomaly detection and root-cause analysis a plus
  • Strong communication skills, with the ability to translate technical reliability data into insight for non-technical stakeholders
  • A coaching mindset, with genuine enthusiasm for developing engineers and building a strong team culture
What’sthe Score?
WhyOpenBet?
  • The Playground:Join a team of innovators, disruptors, andgame-changerswho are reshaping the future of betting and gaming.
  • The Mission:Be part of a mission-driven organizationthat'scommitted to revolutionizing the way the world plays.
  • The Impact:Make a real impact on the world stage, leavinga lasting legacythat transcends boundaries and inspires generations to come.
  • The Culture:Immerse yourself in a culture of creativity, collaboration, and curiosity, where every idea is welcomed, every voice is heard, and every dream is encouraged.
  • The Future:Join us on the journey to build the future of betting and gaming, one game-changing innovation at a time.
What we can offer YOU:
  • Attractive benefits, an open and supportive environment as well as a modern and exciting workplace
  • The opportunity to interact with global teams on a regular basis as you and our business continues to develop & grow
  • Tangible and genuine development - atOpenBet, you can take your career where you want it to go!
  • And ifthat’snot enough;enjoyflexibleworkingwhilst we provide you with theguidanceanddevelopmentskillsyou need to progress andenhance your career
  • We have a collaborative office environment with our team members in office 2 days per week.

AtOpenBet, we celebrate diversity and believe in creating an inclusive environment where every voice is valued and respected.We'recommitted to building a team that reflects the rich tapestry of humanity, embracing individuals from allwalks of life, backgrounds, and identities. Join us in shaping the future of iGaming, where diversityisn'tjust celebrated— it'scelebrated.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

OpenBet • Kentucky

On-site
USD 90,000 - 130,000
Market leading benefits
Career development
Work with global teams
Operations Support Engineer
Operations Support Engineer

OpenBet • Kentucky

Hybrid
USD 46,000 - 70,000
Flexible working
Attractive benefits
Modern, exciting workplace
Software Engineering Manager I
Software Engineering Manager I

bet365 • Denver (CO)

Hybrid
USD 110,000 - 160,000
SRE Lead: Observability & Reliability at Scale
SRE Lead: Observability & Reliability at Scale

OpenBet • Kentucky

Hybrid
USD 140,000 - 200,000
Competitive benefits
Global collaboration
Flexible working
Senior Back-End Engineer, Sports Platform
Senior Back-End Engineer, Sports Platform

bet365 Group • Denver (CO)

Hybrid
USD 125,000 - 155,000
Health & Wellness plans
PTO 33 days
401(k) with company match
+1
Sr. Site Reliability Engineer (SRE)
Sr. Site Reliability Engineer (SRE)

Scientific Games, LLC • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Raleigh (NC)

On-site
USD 120,000 - 180,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 170,000 - 210,000
Healthcare
Retirement planning
Volunteer days
+1
DevOps Engineer - Player Account Management
DevOps Engineer - Player Account Management

Betsson Group • Kentucky

Hybrid
USD 90,000 - 150,000
Lunch allowance
Private life insurance
Team building budget
+5
Backend Engineering Tech Lead
Backend Engineering Tech Lead

Betsson Group • Kentucky

On-site
USD 110,000 - 170,000
Monthly Lunch Allowance
Private & Life Insurance for you and a
Team Building Budget
+2