Site Reliability Engineer — AI-Driven Reliability (Hybrid)

Hackajob Ltd

West Midlands

Hybrid

GBP 90,000 - 130,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

bet365 is seeking a Site Reliability Engineer to shape the stability of the systems behind every click and live change. You will join a team that protects and improves availability, performance and resilience across a complex technical estate, blending software engineering with automation and incident response.

You will work across SRE, development and IT Operations, embedding reliability throughout the software lifecycle, leading technical work and sharing knowledge to lift standards within the

Qualifications

  • Software engineering background with Python, Golang, JavaScript or similar language.
  • Knowledge of modern development practices, including testing, source control and delivery lifecycles.
  • An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management.
  • Hands‑on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty.
  • Proficiency in shell scripting for automation and system management.
  • Experience with Infrastructure as Code, including Terraform and Ansible.
  • Knowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection and traffic management.
  • Ability to troubleshoot distributed systems across edge, network, platform, application, dependency and origin layers.
  • Experience working in a large‑scale, 24/7 enterprise where uptime, performance and stability are critical.
  • Practical experience using LLM platforms and coding assistants safely to improve productivity, quality and root‑cause analysis.

Responsibilities

  • Develop and maintain resilient tools, operational APIs and automation for effective system management.
  • Use orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.
  • Write and contribute to code, telemetry and instrumentation that improve service reliability and observability.
  • Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.
  • Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
  • Diagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.
  • Participate in live incident response, post-mortems and root‑cause analysis to prevent recurrence.
  • Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.
  • Drive initiatives that improve reliability, observability, performance and continuous improvement across teams.
  • Mentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.

Skills

Python
Golang
JavaScript
Testing
Source control
SRE principles
Observability tools
Shell scripting
Terraform
Ansible
Edge platform
LLM platforms
Incident management
Distributed systems

Tools

OpenTelemetry
Grafana
Splunk
New Relic
PagerDuty

Job description

bet365 is seeking a Site Reliability Engineer to shape the stability of the systems behind every click and live change. You will join a team that protects and improves availability, performance and resilience across a complex technical estate, blending software engineering with automation and incident response.

You will work across SRE, development and IT Operations, embedding reliability throughout the software lifecycle, leading technical work and sharing knowledge to lift standards within the

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — AI-Driven Ops & Observability
Site Reliability Engineer — AI-Driven Ops & Observability

bet365 • Stoke-on-Trent

Hybrid
GBP 65,000 - 110,000
Hybrid AI-Driven SRE & Reliability Engineer
Hybrid AI-Driven SRE & Reliability Engineer

bet365 Group • United Kingdom

On-site
GBP 70,000 - 110,000
Global Site Reliability Engineer (Hybrid)
Global Site Reliability Engineer (Hybrid)

bet365 • Burslem

Hybrid
GBP 70,000 - 110,000
AI-Driven SRE & Software Engineer
AI-Driven SRE & Software Engineer

bet365 Group • Manchester

Hybrid
GBP 100,000 - 130,000
Hybrid work policy
Hybrid SRE: Build Reliable, Scalable Systems
Hybrid SRE: Build Reliable, Scalable Systems

bet365 • Manchester

Hybrid
GBP 70,000 - 110,000
Hybrid work from home policy
Site Reliability Engineer - Hybrid, Observability
Site Reliability Engineer - Hybrid, Observability

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Hybrid SRE: AI-Driven Reliability & Observability
Hybrid SRE: AI-Driven Reliability & Observability

bet365 Group • Manchester

Hybrid
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Site Reliability Engineer
Site Reliability Engineer

bet365 • Stoke-on-Trent

Hybrid
GBP 65,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

bet365 • Burslem

On-site
GBP 70,000 - 110,000
Site Reliability Engineer – Hybrid, Scale & Resilience
Site Reliability Engineer – Hybrid, Scale & Resilience

William Hill PLC • Leeds

On-site
GBP 65,000 - 90,000
Family support
Retail discounts
Competitive salary and pension
+4