Lead Site Reliability Engineer

NICE Systems

Southampton

Hybrid

GBP 40,000 - 60,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NICE Public Safety in the United Kingdom is seeking a senior SRE to join a hybrid Cloud Platform Engineering team. You will guard production, drive reliability improvements, and lead root-cause investigations alongside DevOps and engineering peers.

You will design and implement monitoring, dashboards, and alerting using Grafana, Prometheus, OpenTelemetry, and Azure Monitor. The role emphasizes automation, cost optimization, security, and scalable cloud solutions.

Qualifications

  • 6+ years of Site Reliability Engineering experience.
  • Strong technical, analytical, and troubleshooting skills.
  • In-depth knowledge of databases and data handling (MS-SQL, Elasticsearch, YML, JSON, XML).
  • Experience with Azure cloud.
  • Experience with programming or advanced scripting (Python, PowerShell, or C#).
  • Experience with infrastructure as code and version control (ARM, BICEP, Git).
  • Experience with monitoring/alerting/dashboarding (Azure Monitor, Prometheus, Grafana, Elasticsearch).
  • Experience supporting live cloud services and platforms.
  • Developing queries for dashboards/alerts for microservices and custom metrics.
  • Kubernetes and containerization experience, ideally AKS.
  • Exposure to Azure DevOps CI/CD pipelines.
  • Strong design and implementation strategies for infrastructure as code.
  • AI tools to automate work is a plus.
  • Strong communication with customers and internal teams.
  • Methodical troubleshooting and attention to detail.
  • Multitasking and time-management skills.
  • Significant experience in platform engineering, observability, and provisioning.
  • Ability to translate business needs into scalable cloud solutions.
  • Clear communication with non-security SMEs; knowledge of ISO 27001, Cyber Essentials+, FEDRAMP is a plus.
  • UK residency 5 years and eligible for NPPV3 security clearance.

Responsibilities

  • Act as part of a team of SREs to guard production and manage reliability backlog and improvements.
  • Lead investigations into outages, performance, and cost issues.
  • Automate low-value tasks while balancing project delivery demands.
  • Provide technical leadership to Cloud Operations and Support teams.
  • Collaborate with DevOps/engineering to establish SLOs, SLAs, and error budgets.
  • Develop and configure monitoring dashboards and alerts in Grafana and Azure Monitor.
  • Install and configure observability platforms (Grafana, Prometheus, Azure Monitor, OpenTelemetry).
  • Develop Bicep modules for monitoring infrastructure and deploy them.
  • Optimize system performance, cost, and security through regular reviews and tuning.

Skills

SRE experience
Analytical skills
Troubleshooting
Communication with customers
Questioning to confirm understanding
Platform engineering
Observability
Cloud platforms

Tools

Azure
C#
CI/CD
Git
Grafana
OpenTelemetry
Terraform
Prometheus
Kubernetes
MS-SQL
ElasticSearch
PowerShell
Python
JSON
XML
ARM
BICEP

Job description

Salary: £40,000 - 60,000 per year

Requirements:
  • We are looking for someone with 6+ years of experience in Site Reliability Engineering.
  • We need strong technical, analytical, and troubleshooting skills.
  • We value in-depth knowledge of databases and data handling, including MS-SQL, Elasticsearch, YML, JSON, and XML.
  • We require experience with Azure cloud.
  • We need significant experience in programming or advanced scripting such as Python, PowerShell, or C#.
  • We require experience with infrastructure/configuration as code and version control, including ARM, BICEP, and Git.
  • We need strong experience managing monitoring, alerting, and dashboarding platforms such as Azure Monitor, Prometheus, Grafana, and Elasticsearch.
  • We look for demonstrable experience supporting live cloud services and platforms.
  • We need expertise in developing queries for dashboards and alerting for microservices.
  • We need expertise in developing custom metrics for microservices.
  • We require experience with Kubernetes and containerization, ideally AKS.
  • We value exposure to Azure DevOps pipelines and CI/CD.
  • We need strong experience in infrastructure as code, including design and implementation strategies.
  • Experience with AI tools to automate and accelerate work is a plus.
  • We need efficient, effective, and respectful communication skills with customers and internal teams.
  • We need a good listener who can identify and validate assumptions.
  • We need someone able to use effective questioning to confirm understanding of customer problems and provide help to solve them.
  • We need methodical troubleshooting, technical skill, and attention to detail to diagnose problems and reproduce issues in a local environment.
  • We need strong multitasking and time-management skills to prioritise and switch between varied tasks.
  • We require significant experience in platform engineering, observability, and provisioning.
  • We need proven ability to develop and implement a strategic vision for platform services, observability, and provisioning.
  • We need a strong understanding of cyber security principles, governance, and compliance frameworks.
  • We need a strong understanding and experience of cloud platforms, containerisation, and microservices architecture.
  • We need a broad background across information technology with the ability to communicate clearly with non-security technical SMEs.
  • We need strong proficiency in technical scoping, architecture design, and integration of security tools and processes.
  • We need the ability to translate business needs into scalable, user-centric cloud solutions.
  • We need excellent communication and collaboration skills, with a focus on thought leadership and solution development.
  • We value experience in both operational and transformation roles, or a clear working understanding of both perspectives.
  • We value knowledge of compliance with relevant frameworks, including ISO 27001, Cyber Essentials +, or FEDRAMP.
  • The successful candidate must have lived in the UK for 5 years and be eligible to obtain NPPV3 + Security Clearance.
Responsibilities:
  • We act as part of a team of SREs that serve as the gatekeepers of production and actively manage the work backlog and reliability improvements.
  • We lead investigations into root cause outages, performance issues, and cost issues.
  • We lead initiatives to automate low-value tasks while balancing project delivery demands.
  • We provide technical leadership to wider Cloud Operations and Support teams and oversee the products and services they support.
  • We collaborate with DevOps and engineering teams to establish and enforce SLOs, SLAs, and error budgets.
  • We develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor.
  • We install and configure the observability platform, including tools like Grafana, Prometheus, Azure Monitor, and OpenTelemetry.
  • We develop Bicep modules for monitoring infrastructure and deploy them.
  • We optimize system performance, cost, and security through regular reviews and tuning.
Technologies:
  • AI
  • ARM
  • Azure
  • C#
  • CI/CD
  • Cloud
  • DevOps
  • ElasticSearch
  • Git
  • Grafana
  • Support
  • JSON
  • Kubernetes
  • MS-SQL
  • OpenTelemetry
  • PowerShell
  • Prometheus
  • Python
  • SQL
  • Security
  • XML
  • microservices
  • Terraform
More:

We are NICE Public Safety, providing state-of-the-art software as a service solutions for the Public Safety & Justice market, including multi-media evidence management and Emergency Contact Centres for a worldwide customer base. We are expanding our Cloud Platform Engineering team to deliver exemplary service and ensure our cloud platforms remain observable, measurable, reliable, scalable, and maintainable. This is a hands-on, individual contributor role reporting into our Director, Engineering. NiCE is a global innovation powerhouse with more than 8,500 employees across 30+ countries, serving 25,000+ businesses including 85 of the Fortune 100. We are recognized for our leadership in AI, cloud, and digital, and we are proud to be an equal opportunity employer. The role is hybrid.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

NICE • Southampton

On-site
GBP 60,000 - 80,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Nice • Greater London

Hybrid
GBP 70,000 - 110,000
NICE-FLEX hybrid model
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
Cloud Operations Service Reliability Engineer
Cloud Operations Service Reliability Engineer

A&O Shearman • Carrickfergus

On-site
GBP 45,000 - 73,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VIQU IT Recruitment • Kingston

On-site
GBP 68,000 - 83,000
On-call allowance
Bonus
Lead Site Reliability Engineer - Glasgow
Lead Site Reliability Engineer - Glasgow

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer - Edinburgh
Lead Site Reliability Engineer - Edinburgh

Inspire People • City of Edinburgh

Hybrid
GBP 72,000 - 88,000
Site Reliability Engineer - NS London
Site Reliability Engineer - NS London

BAE Systems Digital Intelligence • Greater London

Hybrid
GBP 50,000 - 70,000
Hybrid working environment
On-call allowances
Overtime benefits for night shifts
Site Reliability Engineer – NS London
Site Reliability Engineer – NS London

BAE Systems • Greater London

Hybrid
GBP 45,000 - 70,000
Hybrid working flexibility
On-call allowances
Overtime benefits
Site Reliability Engineer (Sheffield)
Site Reliability Engineer (Sheffield)

Caspian One • Sheffield

Hybrid
GBP 74,000 - 96,000