Software Engineer, Site Reliability Engineering

bet365 Group

Stoke-on-Trent

On-site

GBP 60,000 - 80,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Eye care
Flu Vaccinations
Life Assurance

Job summary

bet365 Group, located in Stoke-on-Trent, is hiring a Site Reliability Engineer. This role focuses on enhancing system reliability and observability through strong engineering practices. Key responsibilities include developing tools, automating processes, and ensuring effective incident resolution.

The ideal candidate should have comprehensive knowledge of software development, observability tools like Splunk and Grafana, and experience with automation practices. This position offers hybrid working arrangements and various perks including eye care and life assurance.

Qualifications

  • Strong software engineering skills with a focus on reliability.
  • Experience in automating tasks using scripting.
  • Hands-on experience with AI tools and LLM platforms.

Responsibilities

  • Enhance system reliability and observability through engineering.
  • Develop tools for effective system management.
  • Collaborate across functions to embed best practices.

Skills

Knowledge of modern software development techniques
Excellent knowledge of SRE principles
Proficiency in shell scripting
Experience with IaC tools like Ansible and Terraform
Knowledge of observability tools like Splunk and Grafana
AI-native engineering approach

Tools

OpenTelemetry
Grafana
Splunk
New Relic
PagerDuty
Terraform
Ansible

Job description

As a Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices.

Full-time

Closes 05/08/2026

You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing for, and progressing the performance and availability of our critical systems.

Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices.

You will leverage AI tools and LLM platforms in your daily work to reduce toil, drive autonomous operations, and optimise system health, while engineering automation and tooling for effective service management.

Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement.

This role is eligible for inclusion in the Company's hybrid working from home policy.

Preferred Skills and Experience
  • Knowledge and experience of modern software development techniques and lifecycles.
  • Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction.
  • Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty.
  • Proficiency in shell scripting for automation and system management tasks.
  • Experience with Infrastructure as Code (IaC), automation and orchestration tools such as Ansible and Terraform.
  • Prior experience working in a large scale, 24/7 enterprise where system uptime and stability is of paramount importance to the business.
  • An AI-native engineering approach, with hands‑on experience using LLM platforms and coding assistants to improve productivity and quality, and the ability to integrate AI‑driven telemetry for advanced observability, predictive insights and root‑cause analysis.
What you will be doing
  • Developing and maintaining tools that facilitate effective management of our systems, ensuring they are operationally efficient and resilient.
  • Working with automation and orchestration platforms to automate manual activity and reduce toil.
  • Writing and contributing to code that enhances the reliability and observability of services, including telemetry, operational APIs and tooling.
  • Building sophisticated dashboards using a range of telemetry data and dash boarding technologies like Grafana, Splunk and New Relic.
  • Actively participating in live incident resolution and post‑mortem analysis, providing effective remediation strategies to improve overall system health and prevent future issues.
  • Driving initiatives to enhance system reliability and observability, contributing to a culture of continuous improvement.
  • Maintaining and administering existing monitoring and analytic toolsets.
  • Mentoring colleagues in use of new technologies or practices.
  • Working with IT Operations to provide and support the use of critical tooling that will enable increasing levels of value to the Business.
Bonus
  • Eye care and Flu Vaccinations
  • Life Assurance
Life at bet365

We are a unique global operator with passion and drive to be the best in the industry. Our values form the foundation of culture and shape the unique way that we work. People are our superpower and we support you to be the best you can be.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

bet365 Group • Manchester

On-site
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Stoke-on-Trent

Hybrid
GBP 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Manchester

Hybrid
GBP 60,000 - 80,000
Software Developer, Trading and Tools
Software Developer, Trading and Tools

bet365 Group • Manchester

Hybrid
GBP 40,000 - 60,000
Eye care
Flu Vaccinations
Life Assurance
Software Developer, Trading and Tools
Software Developer, Trading and Tools

bet365 Group • Stoke-on-Trent

Hybrid
GBP 40,000 - 60,000
Eye care
Flu vaccinations
Life assurance
Technical Lead, Risk & Regulatory
Technical Lead, Risk & Regulatory

bet365 Group • Manchester

Hybrid
GBP 90,000 - 130,000
Eye care and Flu Vaccinations
Life Assurance
Senior Infrastructure Automation Engineer
Senior Infrastructure Automation Engineer

bet365 Group • Manchester

Hybrid
GBP 70,000 - 110,000
Eye care
Flu Vaccinations
Life Assurance
Business Systems Software Developer
Business Systems Software Developer

bet365 Group • Stoke-on-Trent

Hybrid
GBP 50,000 - 70,000
Eye care and Flu Vaccinations
Life Assurance
Software Developer, Risk and Regulatory
Software Developer, Risk and Regulatory

bet365 Group • Stoke-on-Trent

Hybrid
GBP 40,000 - 60,000
Eye care
Flu Vaccinations
Life Assurance
Business Systems Software Developer
Business Systems Software Developer

bet365 Group • Manchester

Hybrid
GBP 55,000 - 75,000
Eye care and Flu Vaccinations
Life Assurance