Site Reliability Engineer — AI-Driven Ops & Observability

bet365

Stoke-on-Trent

Hybrid

GBP 65,000 - 110,000

Full time

30 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

bet365 is recruiting a Site Reliability Engineer to shape the stability of systems behind every click, query and live change. The role focuses on availability, performance and resilience, blending software engineering, automation and incident response to reduce toil and strengthen service health across a complex estate.

You will work with OpenTelemetry and observability tooling, contribute to AI-assisted operations and help embed reliability throughout the software development lifecycle across

Qualifications

  • Experience in SRE practices, including incident response and reliability metrics.
  • Hands-on with observability tools (OpenTelemetry, Splunk, New Relic, Grafana).
  • Proficient shell scripting for automation.
  • Experience with Infrastructure as Code (Terraform, Ansible).
  • Familiarity with Cloudflare edge services (DNS, CDN, WAF, DDoS).
  • Ability to troubleshoot distributed systems across edge, network and origin layers.
  • Experience in a large-scale 24/7 enterprise with uptime criticality.
  • Experience using LLM platforms and coding assistants to boost productivity.

Responsibilities

  • Develop and maintain resilient tools, operational APIs and automation for effective system management.
  • Use orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.
  • Write and contribute to code, telemetry and instrumentation that improve service reliability and observability.
  • Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.
  • Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
  • Diagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.
  • Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence.
  • Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.
  • Drive initiatives that improve reliability, observability, performance and continuous improvement across teams.
  • Mentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.

Skills

SLIs & SLOs
Incident management
Observability
Automation
Distributed systems
LLM tooling
Shell scripting

Tools

OpenTelemetry
Splunk
New Relic
Grafana
PagerDuty
Terraform
Ansible
Cloudflare

Job description

bet365 is recruiting a Site Reliability Engineer to shape the stability of systems behind every click, query and live change. The role focuses on availability, performance and resilience, blending software engineering, automation and incident response to reduce toil and strengthen service health across a complex estate.

You will work with OpenTelemetry and observability tooling, contribute to AI-assisted operations and help embed reliability throughout the software development lifecycle across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-Driven SRE: Reliability, Observability & Automation
AI-Driven SRE: Reliability, Observability & Automation

Bet365 • Manchester

Hybrid
GBP 90,000 - 120,000
Hybrid work policy
AI-Driven SRE & Software Engineer
AI-Driven SRE & Software Engineer

bet365 Group • Manchester

Hybrid
GBP 100,000 - 130,000
Hybrid work policy
Site Reliability Engineer - Hybrid, Observability
Site Reliability Engineer - Hybrid, Observability

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Global Site Reliability Engineer (Hybrid)
Global Site Reliability Engineer (Hybrid)

bet365 • Burslem

Hybrid
GBP 70,000 - 110,000
Hybrid AI-Driven SRE & Reliability Engineer
Hybrid AI-Driven SRE & Reliability Engineer

bet365 Group • United Kingdom

On-site
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

bet365 • Stoke-on-Trent

Hybrid
GBP 65,000 - 110,000
Hybrid SRE: AI-Driven Reliability & Observability
Hybrid SRE: AI-Driven Reliability & Observability

bet365 Group • Manchester

Hybrid
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Hybrid SRE: Build Reliable, Scalable Systems
Hybrid SRE: Build Reliable, Scalable Systems

bet365 • Manchester

Hybrid
GBP 70,000 - 110,000
Hybrid work from home policy
Site Reliability Engineer
Site Reliability Engineer

bet365 Group • Manchester

On-site
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Software Engineer, SRE
Software Engineer, SRE

bet365 Group • Manchester

Hybrid
GBP 100,000 - 130,000
Hybrid work policy