AI-Driven SRE: Reliability, Observability & Automation

Bet365

Manchester

Hybrid

GBP 90,000 - 120,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid work policy

Job summary

bet365 is seeking a Site Reliability Engineer to enhance system reliability, observability and performance through an engineering-led approach. You will tackle incidents, improve logging, instrumentation with OpenTelemetry, and drive best practices across the SDLC.

You will apply AI-native engineering, integrate LLM platforms for productivity and telemetry, and collaborate across functions to embed reliability in the software lifecycle. The role supports bet365's hybrid work policy.

Qualifications

  • Excellent knowledge of programming languages including Python, Golang and JavaScript.
  • Knowledge and experience of modern software development techniques and lifecycles.
  • Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction.
  • Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty.
  • Proficiency in shell scripting for automation and system management tasks.
  • Experience with Infrastructure as Code (IaC), automation and orchestration tools such as Ansible and Terraform.
  • Prior experience working in a large scale, 24/7 enterprise where system uptime and stability is of paramount importance to the business.
  • An AI-native engineering approach, with hands-on experience using LLM platforms and coding assistants to improve productivity and quality, and the ability to integrate AI-driven telemetry for advanced observability, predictive insights and root-cause analysis.

Responsibilities

  • Developing and maintaining tools that facilitate effective management of our systems, ensuring they are operationally efficient and resilient.
  • Working with automation and orchestration platforms to automate manual activity and reduce toil.
  • Writing and contributing to code that enhances the reliability and observability of services, including telemetry, operational APIs and tooling.
  • Building sophisticated dashboards using a range of telemetry data and dash boarding technologies like Grafana, Splunk and New Relic.
  • Actively participating in live incident resolution and post-mortem analysis, providing effective remediation strategies to improve overall system health and prevent future issues.
  • Driving initiatives to enhance system reliability and observability, contributing to a culture of continuous improvement.
  • Maintaining and administering existing monitoring and analytic toolsets.
  • Mentoring colleagues in use of new technologies or practices.
  • Working with IT Operations to provide and support the use of critical tooling that will enable increasing levels of value to the Business.

Skills

Python
Golang
JavaScript

Tools

Splunk
New Relic
Grafana
PagerDuty
Ansible
Terraform

Job description

bet365 is seeking a Site Reliability Engineer to enhance system reliability, observability and performance through an engineering-led approach. You will tackle incidents, improve logging, instrumentation with OpenTelemetry, and drive best practices across the SDLC.

You will apply AI-native engineering, integrate LLM platforms for productivity and telemetry, and collaborate across functions to embed reliability in the software lifecycle. The role supports bet365's hybrid work policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-Driven SRE & Observability Engineer
AI-Driven SRE & Observability Engineer

Bet3651 • Manchester

Hybrid
GBP 90,000 - 120,000
Site Reliability Engineer: AI-Powered Observability
Site Reliability Engineer: AI-Powered Observability

bet365 • Stoke-on-Trent

Hybrid
GBP 70,000 - 100,000
Hybrid working from home
Site Reliability Engineer (Hybrid) — AI Observability
Site Reliability Engineer (Hybrid) — AI Observability

Bet3651 • Stoke-on-Trent

Hybrid
GBP 90,000 - 130,000
Hybrid working policy
Hybrid SRE: AI-Driven Reliability & Observability
Hybrid SRE: AI-Driven Reliability & Observability

bet365 Group • Manchester

Hybrid
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Hybrid AI-Powered Site Reliability Engineer
Hybrid AI-Powered Site Reliability Engineer

JobCubby • Manchester

Hybrid
GBP 90,000 - 120,000
Software Engineer, SRE
Software Engineer, SRE

Bet3651 • Manchester

Hybrid
GBP 90,000 - 120,000
Software Engineer, SRE
Software Engineer, SRE

Bet3651 • Stoke-on-Trent

Hybrid
GBP 90,000 - 130,000
Hybrid working policy
Site Reliability Engineer - Hybrid, Observability
Site Reliability Engineer - Hybrid, Observability

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Hybrid SRE: Build Reliable, Scalable Systems
Hybrid SRE: Build Reliable, Scalable Systems

bet365 • Manchester

Hybrid
GBP 70,000 - 110,000
Hybrid work from home policy
Software Engineer, SRE
Software Engineer, SRE

bet365 • Stoke-on-Trent

Hybrid
GBP 70,000 - 100,000
Hybrid working from home