Senior Site Reliability Engineer: Automation & Observability

National Health Service

Greater London

Hybrid

GBP 42,000 - 52,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid working
Flexible working arrangements
On-site core HQs

Job summary

UK Health Security Agency (UKHSA) is recruiting a permanent Site Reliability Engineer to join the HPC & SRE engineering team, combining software and systems engineering to build, improve, and operate reliable production systems. The role is available full-time, part-time, as a job share, or with flexible working.

We offer hybrid working from our core HQs or scientific campuses with 60% on site. Salary is £41,983–£52,113 per year, with a market pay supplement up to £5,000 pro rata, subject to

Qualifications

  • Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer, or similar role.
  • Proficient in Python, PowerShell, or Bash scripting.
  • Strong understanding of Linux/Unix and Windows systems, networking and distributed systems.
  • Experience with observability tools (Prometheus, Grafana, Datadog) and alerting systems.
  • Familiarity with infrastructure automation tools (Terraform, Ansible, PowerShell, Helm).
  • Excellent communication and collaboration skills; able to respond to unexpected demands.
  • Desirable: CI/CD, cloud platforms (AWS, GCP, Azure), and Kubernetes.

Responsibilities

  • Ensure services are stable, scalable, and automated.
  • Respond to production incidents and conduct root cause analysis and post-incident reviews.
  • Identify bottlenecks, tune performance, and plan capacity for current and future workloads.
  • Contribute to monitoring, alerting, and observability improvements.
  • Develop automation and tooling; use Infrastructure as Code to improve reliability.
  • Define and track SLOs/SLIs and error budgets; drive operational improvements.
  • Promote SRE practices across the organization and integrate reliability into development.
  • Collaborate with software, DevOps, and infrastructure teams to improve deployment workflows.
  • Maintain runbooks, post-incident reports, and provide mentorship to engineers.

Skills

Python
PowerShell
Bash
Linux/Unix
Windows
Observability

Tools

AWS
Azure
GCP
Ansible
Terraform
Helm
Kubernetes
Prometheus
Grafana
Datadog
CI/CD

Job description

UK Health Security Agency (UKHSA) is recruiting a permanent Site Reliability Engineer to join the HPC & SRE engineering team, combining software and systems engineering to build, improve, and operate reliable production systems. The role is available full-time, part-time, as a job share, or with flexible working.

We offer hybrid working from our core HQs or scientific campuses with 60% on site. Salary is £41,983–£52,113 per year, with a market pay supplement up to £5,000 pro rata, subject to

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

National Health Service • Greater London

Hybrid
GBP 42,000 - 52,000
Hybrid working
Flexible working arrangements
On-site core HQs
Senior SRE Engineer: Automation & Observability
Senior SRE Engineer: Automation & Observability

UK Health Security Agency • Liverpool

On-site
GBP 60,000 - 90,000
Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)
Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)

UK Health Security Agency • Liverpool

On-site
GBP 60,000 - 90,000
Senior Site Reliability Engineer (LON)
Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd • Greater London

Hybrid
GBP 90,000 - 150,000
Benefits as Cash
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

Sanderson Government & Defence • Greater London

On-site
GBP 35,000 - 75,000
Flexible salary range reflecting seniority
Hybrid working and startup autonomy
Opportunity to shape platforms and engineering practices
Site Reliability Engineer – NS London
Site Reliability Engineer – NS London

BAE Systems • Greater London

On-site
GBP 45,000 - 70,000
Hybrid working flexibility
On-call allowances
Overtime benefits
Site Reliability Engineer - NS London
Site Reliability Engineer - NS London

BAE Systems Digital Intelligence • Greater London

On-site
GBP 50,000 - 70,000
Hybrid working environment
On-call allowances
Overtime benefits for night shifts
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Site Reliability Engineer - AI-Driven SaaS Platform
Site Reliability Engineer - AI-Driven SaaS Platform

Obsidian Security • Salford

On-site
GBP 85,000 - 103,000
Competitive compensation with equity
Comprehensive healthcare
Flexible paid time off
+2
Senior Site Reliability Engineer — Cloud & Observability
Senior Site Reliability Engineer — Cloud & Observability

GCA Altium • Cambridge

On-site
GBP 90,000 - 140,000