Senior Site Reliability Engineer (SRE)

National Health Service

Greater London

Hybrid

GBP 42,000 - 52,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Hybrid working
Flexible working arrangements
On-site core HQs

Job summary

UK Health Security Agency (UKHSA) is recruiting a permanent Site Reliability Engineer to join the HPC & SRE engineering team, combining software and systems engineering to build, improve, and operate reliable production systems. The role is available full-time, part-time, as a job share, or with flexible working.

We offer hybrid working from our core HQs or scientific campuses with 60% on site. Salary is £41,983–£52,113 per year, with a market pay supplement up to £5,000 pro rata, subject to

Qualifications

  • Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer, or similar role.
  • Proficient in Python, PowerShell, or Bash scripting.
  • Strong understanding of Linux/Unix and Windows systems, networking and distributed systems.
  • Experience with observability tools (Prometheus, Grafana, Datadog) and alerting systems.
  • Familiarity with infrastructure automation tools (Terraform, Ansible, PowerShell, Helm).
  • Excellent communication and collaboration skills; able to respond to unexpected demands.
  • Desirable: CI/CD, cloud platforms (AWS, GCP, Azure), and Kubernetes.

Responsibilities

  • Ensure services are stable, scalable, and automated.
  • Respond to production incidents and conduct root cause analysis and post-incident reviews.
  • Identify bottlenecks, tune performance, and plan capacity for current and future workloads.
  • Contribute to monitoring, alerting, and observability improvements.
  • Develop automation and tooling; use Infrastructure as Code to improve reliability.
  • Define and track SLOs/SLIs and error budgets; drive operational improvements.
  • Promote SRE practices across the organization and integrate reliability into development.
  • Collaborate with software, DevOps, and infrastructure teams to improve deployment workflows.
  • Maintain runbooks, post-incident reports, and provide mentorship to engineers.

Skills

Python
PowerShell
Bash
Linux/Unix
Windows
Observability

Tools

AWS
Azure
GCP
Ansible
Terraform
Helm
Kubernetes
Prometheus
Grafana
Datadog
CI/CD

Job description

Salary: £61,000 - 101,000 per year

Requirements:
  • We require experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer, or in a similar role.
  • We require coding skills in programming or scripting languages such as Python, PowerShell, or Bash.
  • We require an understanding of Linux/Unix and Windows systems, networking, and distributed systems.
  • We require experience with observability tools such as Prometheus, Grafana, or Datadog, and alerting systems.
  • We require an understanding of infrastructure automation tools such as Terraform, Ansible, PowerShell, or Helm.
  • We require excellent communication and collaboration skills, strong problem-solving ability, and the ability to respond to unexpected demands.
  • Desirable: experience with CI/CD pipelines, cloud platforms such as AWS, GCP, or Azure, and container orchestration such as Kubernetes.
  • Desirable: experience with post-incident reviews, promoting SRE practices across an organization, or training and mentoring junior engineers.
  • We assess candidates on changing and improving, working together, managing a quality service, and delivering at pace; the application requires a supporting statement of up to 1,000 words and a presentation on automating a complex operational process.
  • Successful candidates must pass a basic Disclosure and Barring Service check and meet Security Check clearance requirements. The posting states that candidates would normally have been resident in the UK for the last five years for these checks.
Responsibilities:
  • We expect you to help ensure our services are stable, scalable, performant, and automated.
  • Respond to production incidents, troubleshoot issues, restore services quickly, and conduct root cause analysis and post-incident reviews.
  • Identify system bottlenecks, tune performance, and plan capacity for current and future workloads.
  • Contribute to effective monitoring, alerting, and observability, refining practices to identify issues early and improve response times.
  • Develop automation and tooling to reduce repetitive manual work and operational overhead; write clear, maintainable, well-tested code and use Infrastructure as Code to improve reliability.
  • Contribute to defining, tracking, and improving SLOs, SLIs, and error budgets, and prioritize operational improvements.
  • Promote SRE principles and work with stakeholders to integrate reliability practices into the development lifecycle.
  • Collaborate with software engineering, DevOps, and infrastructure teams to improve deployment and operational workflows and encourage shared responsibility for reliability.
  • Maintain technical documentation, runbooks, and post-incident reports, and provide training and mentorship to engineering teams.
Technologies:
  • AWS
  • Ansible
  • Azure
  • Bash
  • CI/CD
  • Cloud
  • Datadog
  • DevOps
  • GCP
  • Grafana
  • Helm
  • Kubernetes
  • Linux
  • PowerShell
  • Prometheus
  • Python
  • Security
  • Terraform
  • Unix
  • Windows
  • Support
  • LESS
More:

We are the UK Health Security Agency (UKHSA), and we value an inclusive workplace where everyone matters and differences help us develop innovative solutions. We are recruiting a permanent Site Reliability Engineer to join our HPC & SRE engineering team, combining software and systems engineering to build, improve, and operate reliable production systems. The role is available full-time, part-time, as a job share, or with flexible working. We offer hybrid working from our core HQs or scientific campuses in Birmingham, Chilton, Leeds, Liverpool, London, or Porton; the role normally requires at least 60% of contractual hours on site, averaged over a month. Salary is £41,983–£52,113 per year, depending on location, with a market pay supplement of up to £5,000 pro rata, subject to review.

last updated 40 week of 2026

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer: Automation & Observability
Senior Site Reliability Engineer: Automation & Observability

National Health Service • Greater London

Hybrid
GBP 42,000 - 52,000
Hybrid working
Flexible working arrangements
On-site core HQs
Site Reliability Engineer - NS London
Site Reliability Engineer - NS London

BAE Systems Digital Intelligence • Greater London

On-site
GBP 50,000 - 70,000
Hybrid working environment
On-call allowances
Overtime benefits for night shifts
Site Reliability Engineer – NS London
Site Reliability Engineer – NS London

BAE Systems • Greater London

On-site
GBP 45,000 - 70,000
Hybrid working flexibility
On-call allowances
Overtime benefits
Senior SRE (AWS)
Senior SRE (AWS)

VIQU IT Recruitment • Kingston

On-site
GBP 68,000 - 83,000
Bonus
On-call allowance
SRE Technical Lead
SRE Technical Lead

83zero Ltd • United Kingdom

Hybrid
GBP 90,000 - 110,000
Salary up to 100,000
5% annual bonus
Hybrid working model
+1
Senior Site Reliability Engineer (LON)
Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd • Greater London

Hybrid
GBP 90,000 - 150,000
Benefits as Cash
Hybrid work model
Lead Site Reliability Engineer - Edinburgh
Lead Site Reliability Engineer - Edinburgh

Inspire People • City of Edinburgh

On-site
GBP 72,000 - 88,000
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Lead SRE - Charing Cross
Lead SRE - Charing Cross

Hackajob Ltd • Leeds

On-site
GBP 61,000 - 101,000
DevOps / Site Reliability Engineer (SRE)
DevOps / Site Reliability Engineer (SRE)

SCC • United Kingdom

Hybrid
GBP 70,000 - 110,000