Senior SRE: Azure Cloud Platform & Observability Leader

Spectrum IT Recruitment

Southampton

Hybrid

GBP 80,000 - 110,000

Full time

40 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Spectrum IT Recruitment seeks a Senior Site Reliability Engineer to join our client in Southampton, a hybrid role focused on ensuring Azure-based cloud platforms are observable, scalable and secure. You will drive automation, mentoring across Cloud Operations, and cost-efficient improvements.

You’ll bring expert Azure and Kubernetes skills, plus hands-on experience with monitoring stacks, IaC, and incident leadership in live-service environments.

Qualifications

  • At least six years' commercial experience within Site Reliability Engineering or a closely related cloud platform role.
  • Demonstrable experience supporting business-critical cloud platforms and live production services.
  • Strong hands-on knowledge of Microsoft Azure.
  • Production experience with Kubernetes and containerised workloads, ideally using AKS.
  • Extensive experience in platform engineering, cloud provisioning and observability.
  • Strong monitoring, alerting and dashboarding experience using technologies such as: Azure Monitor, Grafana, Prometheus, OpenTelemetry, Elasticsearch.
  • Experience creating custom metrics, queries, dashboards and alerts for microservices.
  • Advanced scripting or software development skills using PowerShell, Python, C# or a comparable language.
  • Strong Infrastructure as Code experience using Bicep, ARM or Terraform.
  • Experience using Git or another version-control platform.
  • Good knowledge of Microsoft SQL Server, Elasticsearch and structured data formats including YAML, JSON and XML.
  • Strong understanding of microservices architecture, cloud platforms and containerisation.
  • Experience defining or working with SLOs, SLAs, SLIs and error budgets.
  • Excellent troubleshooting and root-cause analysis skills.
  • Experience designing scalable, secure and maintainable cloud solutions.
  • Strong understanding of cybersecurity principles, governance and compliance.
  • Experience operating across both transformation projects and live-service environments.

Responsibilities

  • Work as part of the Site Reliability Engineering team responsible for protecting and improving production environments.
  • Manage and prioritise a technical backlog of reliability, scalability and operational improvements.
  • Lead investigations into service outages, performance degradation, platform reliability and cloud expenditure.
  • Conduct detailed root-cause analysis and ensure corrective actions are implemented.
  • Identify repetitive operational activities and replace them with sustainable automation.
  • Provide technical leadership and guidance to Cloud Operations, Support, DevOps and Engineering teams.
  • Establish and maintain service level objectives, service level agreements, service level indicators and error budgets.
  • Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices.
  • Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry.
  • Develop custom application and platform metrics to improve operational visibility.
  • Create advanced queries, dashboards and alerts for distributed microservices.
  • Develop reusable Bicep or Terraform modules for monitoring and cloud infrastructure.
  • Support and improve production Kubernetes environments, particularly Azure Kubernetes Service.
  • Review and optimise platform performance, availability, security and cost.
  • Contribute to cloud architecture, technical scoping and the implementation of scalable platform solutions.
  • Support continuous improvement across deployment, provisioning and operational processes.
  • Help ensure platforms and working practices meet relevant security, governance and compliance requirements.
  • Explore opportunities to use AI-assisted tools to improve automation, troubleshooting and engineering productivity.

Skills

Azure engineering
Kubernetes & AKS
IaC (Bicep/Terraform)
Observability & monitoring
SRE/DevOps mindset

Tools

Azure Monitor
Grafana
Prometheus
OpenTelemetry
Kubernetes
AKS
Terraform
Bicep

Job description

Spectrum IT Recruitment seeks a Senior Site Reliability Engineer to join our client in Southampton, a hybrid role focused on ensuring Azure-based cloud platforms are observable, scalable and secure. You will drive automation, mentoring across Cloud Operations, and cost-efficient improvements.

You’ll bring expert Azure and Kubernetes skills, plus hands-on experience with monitoring stacks, IaC, and incident leadership in live-service environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Azure SRE — Platform Reliability & Observability
Senior Azure SRE — Platform Reliability & Observability

Spectrum IT • Southampton

Hybrid
GBP 90,000 - 120,000
Remote Senior Azure SRE - Reliability & Automation
Remote Senior Azure SRE - Reliability & Automation

OMEGA, Inc. • United Kingdom

Remote
GBP 100,000 - 140,000
Competitive salary and benefits
Professional development
Certification support
+1
Senior SRE – Hybrid, Cloud & Observability
Senior SRE – Hybrid, Cloud & Observability

Spectrum IT Recruitment • Southampton

Hybrid
GBP 55,000 - 90,000
Life Insurance
Private Medical Insurance
Employee Assistance Programme
+3
Senior Cloud Platform SRE & Observability Lead
Senior Cloud Platform SRE & Observability Lead

NICE • Southampton

On-site
GBP 60,000 - 80,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 80,000 - 110,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Spectrum IT • Southampton

Hybrid
GBP 90,000 - 120,000
Lead Azure SRE: Reliability, Observability & Automation
Lead Azure SRE: Reliability, Observability & Automation

The Nottingham • Nottingham

Hybrid
GBP 85,000 - 120,000
Competitive package
Health & wellbeing resources
35-hour week
+5
Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1
Senior SRE: Cloud Platform Reliability & Observability
Senior SRE: Cloud Platform Reliability & Observability

Renesas Electronics Corporation • Cambridge

Hybrid
GBP 90,000 - 120,000