Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd

Greater London

Hybrid

GBP 90,000 - 150,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Benefits as Cash
Hybrid work model

Job summary

McNally Recruitment Ltd in London seeks a Senior Site Reliability Engineer. The role is largely remote with one day per week in the London office, focused on improving availability, performance, and incident response for cloud-native services.

You will work with feature teams to meet service level objectives, define error budgets, shape release processes, scale systems through automation, and mentor colleagues while engaging stakeholders.

Qualifications

  • At least 10 years hands-on experience including as a Senior SRE
  • Experience with cloud-native microservices, Kubernetes, API management
  • Hands-on with Azure, IaC, PowerShell, JSON, Azure Bicep, ARM and Azure DevOps
  • Terraform experience is essential, moving from Bicep is desirable
  • Experience with Full Stack Observability using Grafana Stack, Log Analytics, AppInsights
  • Strong knowledge of DevOps, IT Service Management and automation with Orchestration and ServiceNow
  • Excellent communication skills and stakeholder engagement

Responsibilities

  • Work closely with feature teams to meet defined service level objectives and continually improve systems and environments
  • Define error budgets to balance risk and reliability
  • Provide structure and support to release processes, suggesting improvements
  • Scale systems sustainably through automation and evolution of processes
  • Coach and guide colleagues and the wider team, leading where required
  • Proactively contribute new ideas to meet short- and long-term goals
  • Balance and manage potential risks
  • Be accountable for day-to-day health of production and non-production environments and respond to incidents
  • Provide input to establish risk tolerance of products and services
  • Communicate incident status updates clearly to teams, customers and stakeholders

Skills

Senior SRE
Kubernetes
Azure
IaC
Terraform
Full Stack Observability
DevOps
ServiceNow automation

Tools

Grafana
Log Analytics
AppInsights
PowerShell
JSON
Azure DevOps
Azure Bicep
ARM
Terraform

Job description

Senior Site Reliability Engineer (London)

We’re working in collaboration to source a Senior Site Reliability Engineer for a large UK client. The role is mostly working remotely, with only 1 day per week being required to work in the London office.

  • In this key role, you’ll improve, drive, and embed non-functional and operational characteristics such as availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning of our products and services
  • You’ll enjoy significant stakeholder interaction, working in collaboration with engineers to ensure a principled approach to deliver change in a safe and secure way
  • This is a chance to join an inclusive team with a collaborative ethos and a commitment to innovation and professional development
  • You’ll work from home some of the time, but you’ll also spend a significant amount of time working from an office or hub
What you’ll do
  • Work closely with our feature team and other colleagues to meet defined service level objectives and continually improve systems and environments.
  • Define error budgets that support finding the right balance between risk and reliability.
  • Provide structure and help to our release process, suggesting and making improvements where possible.
  • Help scale systems sustainably through mechanisms like automation, evolving them by pushing for changes that improve reliability and velocity.
  • Coach and provide guidance to colleagues and the wider team, leading where required.
In addition to this, you’ll:
  • Proactively contribute new ideas and innovations to meet short-term and longer-term goals
  • Continually balance and manage any potential risks
  • Be accountable for the day-to-day health of both production and non-production environments and respond to any incidents as required
  • Provide technical expertise and input to establish the risk tolerance of products and services
  • Communicate incident status updates clearly and frequently to other teams, customers and stakeholders
The skills you’ll need
  • At least 10 years of hands-on experience, including as a Senior SRE with a proactive approach to spotting problems, areas for improvement, and performance bottlenecks.
  • Experience working with cloud-native microservices, including containerisation, management of Kubernetes workloads and API management.
  • Hands-on experience with Azure, Infrastructure as Code (IaC), and technologies such as PowerShell, JSON, Azure Bicep, ARM and Azure DevOps.
  • The client is moving to Terraform, which is essential, moving from Bicep (desirable).
  • Experience with Full Stack Observability using tools such as Grafana Stack, Log Analytics, AppInsights
  • Excellent knowledge of DevOps processes and principles
  • Knowledge of IT Service Management and automation of IT fulfilment processes through Orchestration and ServiceNow
  • Strong communication skills with the ability to proactively engage with a wide range of stakeholders

SALARY INCLUDES 10% Benefits-As-Cash

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 80,000 - 110,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Spectrum IT • Southampton

Hybrid
GBP 90,000 - 120,000
DevOps / Site Reliability Engineer (SRE)
DevOps / Site Reliability Engineer (SRE)

SCC • United Kingdom

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

JAM Recruitment Ltd • City Of London

On-site
GBP 92,000 - 129,000
Site Reliability Engineer - NS London
Site Reliability Engineer - NS London

BAE Systems Digital Intelligence • Greater London

On-site
GBP 50,000 - 70,000
Hybrid working environment
On-call allowances
Overtime benefits for night shifts
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 55,000 - 90,000
Life Insurance
Private Medical Insurance
Employee Assistance Programme
+3
Site Reliability Engineer – NS London
Site Reliability Engineer – NS London

BAE Systems • Greater London

On-site
GBP 45,000 - 70,000
Hybrid working flexibility
On-call allowances
Overtime benefits
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tenth Revolution Group • Knutsford

On-site
GBP 70,000 - 90,000
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

On-site
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14