Software Engineer, SRE

bet365

Stoke-on-Trent

Hybrid

GBP 70,000 - 100,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid working from home

Job summary

bet365 is seeking a Site Reliability Engineer to improve reliability, observability and performance across critical systems. You will apply engineering discipline to protect uptime, instrument services with OpenTelemetry and enhance logging practices.

You will use AI tools and LLM platforms to reduce toil and drive autonomous operations while building automation for effective service management. The role embraces collaboration across functions to embed reliability best practices in the software

Qualifications

  • Proficient in Python, Golang or JavaScript for reliability tooling.
  • Strong understanding of modern software development lifecycles.
  • Deep knowledge of SRE principles and SLI/SLO design.
  • Experience with observability tools such as Splunk, New Relic, Grafana and PagerDuty.
  • Shell scripting for automation and system management tasks.
  • IaC and automation tools like Ansible and Terraform.
  • Experience in large-scale, 24/7 enterprise environments.
  • AI-native engineering with LLM platforms for observability and automation.

Responsibilities

  • Develop and maintain tools to improve system reliability and resilience.
  • Automate manual activities to reduce toil with automation platforms.
  • Contribute code for telemetry and operational APIs to boost observability.
  • Build sophisticated dashboards using Grafana, Splunk and New Relic.
  • Participate in live incident resolution and post-mortem analyses.
  • Drive initiatives to enhance reliability and observability across services.
  • Maintain and administer monitoring and analytics toolsets.
  • Mentor colleagues in new technologies and practices.
  • Collaborate with IT Operations to support critical tooling.

Skills

Python
Golang
JavaScript
Software development
SRE principles
SLIs/SLOs
Observability tools
Shell scripting
Ansible
Terraform
LLM platforms
24/7 operations

Tools

Splunk
New Relic
Grafana
PagerDuty

Job description

Job Description

As a Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices.

You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing for, and progressing the performance and availability of our critical systems.

Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices.

You will leverage AI tools and LLM platforms in your daily work to reduce toil, drive autonomous operations, and optimise system health, while engineering automation and tooling for effective service management.

Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement.

This role is eligible for inclusion in the Company's hybrid working from home policy.

Qualifications
  • Excellent knowledge of programming languages including Python, Golang and JavaScript.
  • Knowledge and experience of modern software development techniques and lifecycles.
  • Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction.
  • Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty.
  • Proficiency in shell scripting for automation and system management tasks.
  • Experience with Infrastructure as Code (IaC), automation and orchestration tools such as Ansible and Terraform.
  • Prior experience working in a large scale, 24/7 enterprise where system uptime and stability is of paramount importance to the business.
  • An AI-native engineering approach, with hands-on experience using LLM platforms and coding assistants to improve productivity and quality, and the ability to integrate AI-driven telemetry for advanced observability, predictive insights and root-cause analysis.
Additional Information
  • Developing and maintaining tools that facilitate effective management of our systems, ensuring they are operationally efficient and resilient.
  • Working with automation and orchestration platforms to automate manual activity and reduce toil.
  • Writing and contributing to code that enhances the reliability and observability of services, including telemetry, operational APIs and tooling.
  • Building sophisticated dashboards using a range of telemetry data and dash boarding technologies like Grafana, Splunk and New Relic.
  • Actively participating in live incident resolution and post-mortem analysis, providing effective remediation strategies to improve overall system health and prevent future issues.
  • Driving initiatives to enhance system reliability and observability, contributing to a culture of continuous improvement.
  • Maintaining and administering existing monitoring and analytic toolsets.
  • Mentoring colleagues in use of new technologies or practices.
  • Working with IT Operations to provide and support the use of critical tooling that will enable increasing levels of value to the Business.

By applying to us you are agreeing to share your Personal Data in accordance with our Recruitment Privacy Notice - https://www.bet365careers.com/privacy-policy

At bet365, we're committed to creating an environment where everyone feels welcome, respected and valued. Where all individuals can grow and develop, regardless of their background. We're Never Ordinary, and we're always striving to be better. If you need any adjustments or accommodations to the recruitment process, at either application or interview, please don't hesitate to reach out.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Sre
Software Engineer, Sre

Bet365 • Manchester

Hybrid
GBP 90,000 - 120,000
Hybrid work policy
Software Engineer, SRE
Software Engineer, SRE

Bet3651 • Manchester

Hybrid
GBP 90,000 - 120,000
Software Engineer, SRE
Software Engineer, SRE

Bet3651 • Stoke-on-Trent

Hybrid
GBP 90,000 - 130,000
Hybrid working policy
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Manchester

Hybrid
GBP 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

bet365 Group • Manchester

On-site
GBP 60,000 - 80,000
Eye care
Flu vaccinations
Life assurance
Software Engineer, SRE
Software Engineer, SRE

JobCubby • Manchester

Hybrid
GBP 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

bet365 • Burslem

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Site Reliability Engineer
Site Reliability Engineer

bet365 • Manchester

Hybrid
GBP 70,000 - 110,000
Hybrid work from home policy
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1