Site Reliability Engineer - Cloud Reliability & AI Observability

BT Group

England

On-site

GBP 70,000 - 110,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BT Group is seeking a Site Reliability Engineering Specialist to independently drive service performance, reliability and availability across digital platforms. You will enable scalable, fault‑tolerant, and cost‑effective cloud services with automation, monitoring, and resilience strategies.

You will mentor other SREs, lead cross‑team engineering discussions, and push for changes that improve reliability and velocity while aligning with BT’s AI/Observability initiatives.

Qualifications

  • A degree in IT, Maths or Science.
  • Deep understanding of full stack monitoring solutions such as Dynatrace to ensure current end to end performance and trends of owned applications.
  • Strong proficiency in one or more programming languages (e.g. Java, Python).
  • Experience with cloud platforms (AWS, Azure, or GCP).
  • Solid understanding of software architecture, design patterns, and microservices.
  • Familiarity with CI/CD tools and DevOps practices.
  • High levels of quality presentation and reporting capabilities to collate output from Managed Service Partners.
  • Resilience to ensure support teams are engaged 24x7x365 to support priority incident resolution.
  • Ability to adapt to latest industry trends
  • Micro-Service functionality
  • Business Process Improvement
  • Growth mindset

Responsibilities

  • Execute implementation of new automation tools, pipelines, and CI/CD practices.
  • Develop platform solutions using AWS cloud, IaC, GitOps, and container technologies.
  • Coordinate a diverse team and deliver testing within time, budget, and quality targets.
  • Automate repeatable tasks to reduce toil and MTTR.
  • Identify and manage risk with proactive controls and mitigations.
  • Lead scale testing to measure and optimize system performance.
  • Perform monitoring analysis to improve stability and security.
  • Design, build, and troubleshoot distributed production systems across on‑prem and cloud environments.
  • Write infrastructure as code to improve availability and efficiency.
  • Implement robust monitoring and run post‑mortems for prevention.
  • Inspect queues and support processing for early issue warning.
  • Lead retrospective actions after major incidents.
  • Analyze complex systems for reliability and resilience.
  • Share SRE best practices and mentor others.
  • Collaborate across teams to deliver improvements aligned with broader initiatives.
  • Mentor other SREs to raise the team’s capabilities.

Skills

IT degree
Dynatrace
Java
Python
Cloud platforms
Microservices architecture
CI/CD tools
DevOps practices
End-to-end performance
Telemetry monitoring
Anomaly detection
SRE fundamentals
Communication skills
Leadership/mentoring

Education

Degree in IT, Maths or Science

Tools

Terraform
Prometheus
Grafana
GitOps

Job description

BT Group is seeking a Site Reliability Engineering Specialist to independently drive service performance, reliability and availability across digital platforms. You will enable scalable, fault‑tolerant, and cost‑effective cloud services with automation, monitoring, and resilience strategies.

You will mentor other SREs, lead cross‑team engineering discussions, and push for changes that improve reliability and velocity while aligning with BT’s AI/Observability initiatives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: Automation & Resilience
Site Reliability Engineer: Automation & Resilience

BT Group • United Kingdom

Hybrid
GBP 70,000 - 110,000
10% on target annual bonus
Online private GP 24/7 for you and 1+1
Paid carers leave up to 2 weeks
+4
Site Reliability Engineer: Resilience & Automation
Site Reliability Engineer: Resilience & Automation

BT Group • Sheffield

On-site
GBP 65,000 - 90,000
10% on target annual bonus
Private GP 24/7 for you and family
Carers leave up to 2 weeks
+1
SRE Engineer: AI-Driven Reliability & Observability(Hybrid)
SRE Engineer: AI-Driven Reliability & Observability(Hybrid)

bet365 • Stoke-on-Trent

Hybrid
GBP 70,000 - 110,000
Hybrid work policy
Hybrid SRE Engineer: AI-Driven Reliability & Observability
Hybrid SRE Engineer: AI-Driven Reliability & Observability

bet365 • Manchester

Hybrid
GBP 75,000 - 110,000
Hybrid working from home policy
SRE Architect: Reliability Leader, Observability
SRE Architect: Reliability Leader, Observability

Hitachi Automotive Systems Americas, Inc. • Greater London

On-site
GBP 95,000 - 130,000
Senior AWS SRE Lead — Cloud Reliability & Automation
Senior AWS SRE Lead — Cloud Reliability & Automation

United States Digital Space LLC • Greater London

Hybrid
GBP 95,000 - 130,000
Hybrid working
Lead Site Reliability Engineer – Resilience & Automation
Lead Site Reliability Engineer – Resilience & Automation

Addition • England

On-site
GBP 90,000 - 120,000
Annual bonus
Pension up to 8.5%
26 days leave + holidays
+5
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Manchester

Hybrid
GBP 60,000 - 80,000
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
SRE Architect: Reliability, Observability & Automation Lead
SRE Architect: Reliability, Observability & Automation Lead

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000