Site Reliability Engineer

bet365

Stoke-on-Trent

Hybrid

GBP 65,000 - 110,000

Full time

25 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

bet365 is recruiting a Site Reliability Engineer to shape the stability of systems behind every click, query and live change. The role focuses on availability, performance and resilience, blending software engineering, automation and incident response to reduce toil and strengthen service health across a complex estate.

You will work with OpenTelemetry and observability tooling, contribute to AI-assisted operations and help embed reliability throughout the software development lifecycle across

Qualifications

  • Experience in SRE practices, including incident response and reliability metrics.
  • Hands-on with observability tools (OpenTelemetry, Splunk, New Relic, Grafana).
  • Proficient shell scripting for automation.
  • Experience with Infrastructure as Code (Terraform, Ansible).
  • Familiarity with Cloudflare edge services (DNS, CDN, WAF, DDoS).
  • Ability to troubleshoot distributed systems across edge, network and origin layers.
  • Experience in a large-scale 24/7 enterprise with uptime criticality.
  • Experience using LLM platforms and coding assistants to boost productivity.

Responsibilities

  • Develop and maintain resilient tools, operational APIs and automation for effective system management.
  • Use orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.
  • Write and contribute to code, telemetry and instrumentation that improve service reliability and observability.
  • Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.
  • Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
  • Diagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.
  • Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence.
  • Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.
  • Drive initiatives that improve reliability, observability, performance and continuous improvement across teams.
  • Mentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.

Skills

SLIs & SLOs
Incident management
Observability
Automation
Distributed systems
LLM tooling
Shell scripting

Tools

OpenTelemetry
Splunk
New Relic
Grafana
PagerDuty
Terraform
Ansible
Cloudflare

Job description

As a Site Reliability Engineer, you will shape the stability of the systems behind every click, query and live change.

Our Site Reliability team protects and improves the availability, performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate.

You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve operational control. The role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight.

Working across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across the wider engineering community.

This role is eligible for inclusion in the company’s hybrid work from home policy.

Qualifications
  • Knowledge of modern development practices, including testing, source control and delivery lifecycles.
  • An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management.
  • Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty.
  • Proficiency in shell scripting for automation and system management.
  • Experience with Infrastructure as Code, including Terraform and Ansible.
  • Knowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection and traffic management.
  • Ability to troubleshoot distributed systems across edge, network, platform, application, dependency and origin layers.
  • Experience working in a large-scale, 24/7 enterprise where uptime, performance and stability are critical.
  • Practical experience using LLM platforms and coding assistants safely to improve productivity, quality and root-cause analysis.
Additional Information
  • Develop and maintain resilient tools, operational APIs and automation for effective system management.
  • Use orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.
  • Write and contribute to code, telemetry and instrumentation that improve service reliability and observability.
  • Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.
  • Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
  • Diagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.
  • Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence.
  • Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.
  • Drive initiatives that improve reliability, observability, performance and continuous improvement across teams.
  • Mentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.

By applying to us you are agreeing to share your Personal Data in accordance with our Recruitment Privacy Notice - https://www.bet365careers.com/privacy-policy

At bet365, we're committed to creating an environment where everyone feels welcome, respected and valued. Where all individuals can grow and develop, regardless of their background. We're Never Ordinary, and we're always striving to be better. If you need any adjustments or accommodations to the recruitment process, at either application or interview, please don’t hesitate to reach out.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Site Reliability Engineer
Site Reliability Engineer

bet365 • Burslem

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

bet365 • Manchester

Hybrid
GBP 70,000 - 110,000
Hybrid work from home policy
Software Engineer, SRE
Software Engineer, SRE

bet365 Group • Manchester

Hybrid
GBP 100,000 - 130,000
Hybrid work policy
Software Engineer, SRE
Software Engineer, SRE

bet365 Group • United Kingdom

Hybrid
GBP 70,000 - 110,000
Software Engineer, SRE
Software Engineer, SRE

bet365 • Manchester

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

bet365 Group • Manchester

On-site
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Site Reliability Engineer - Hybrid, Observability
Site Reliability Engineer - Hybrid, Observability

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Global Site Reliability Engineer (Hybrid)
Global Site Reliability Engineer (Hybrid)

bet365 • Burslem

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer — AI-Driven Ops & Observability
Site Reliability Engineer — AI-Driven Ops & Observability

bet365 • Stoke-on-Trent

Hybrid
GBP 65,000 - 110,000