Service Manager, Site Reliability Engineering

Allstate Northern Ireland Limited

Belfast City District

Hybrid

GBP 75,000 - 95,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Allstate Northern Ireland Limited is seeking a Service Manager - Site Reliability Engineering (SRE) to own health, reliability and observability of critical services. You will lead incident and problem management, ensure governance through ITIL/ITSM, and drive automation and operational maturity across engineering, platform and security teams.

Responsibilities include defining SLAs/SLOs, coordinating major incidents, and guiding change and release processes to ensure production readiness and

Qualifications

  • 4+ years in enterprise technology services, SRE, IT operations, or related.
  • Experience leading incident response and RCA activities in large environments.
  • Strong observability practices with monitoring, alerting, and dashboards.
  • Knowledge of ITIL/ITSM and governance for service management.

Responsibilities

  • Own the health, reliability, availability, and observability of critical applications and services.
  • Establish and report on SLAs, SLOs, KPIs and error budgets.
  • Lead Major Incident Management with cross-functional coordination and RCA outcomes.
  • Govern Change and Release Management and production readiness activities.
  • Drive automation, IaC, and continuous service improvement across the team.

Skills

SRE experience
Leadership & mentoring
Incident management
Observability practices
ITIL/ITSM understanding

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Azure
AWS
GCP
ServiceNow
Terraform/IaC
Monitoring dashboards/logging

Job description

At Allstate, great things happen when our people work together to protect families and their belongings from life’s uncertainties. And for more than 90 years, our innovative drive has kept us a step ahead of our customers’ evolving needs. From advocating for seat belts, air bags and graduated driving laws, to being an industry leader in pricing sophistication, telematics, and, more recently, device and identity protection.**Your role in the team**The Service Manager - Site Reliability Engineering (SRE) is responsible for ensuring the reliability, availability, observability, and operational excellence of technology services while maintaining strong alignment with business objectives. This role serves as the primary owner of service health, incident and problem management, operational governance, continuous service improvement, and stakeholder engagement. The Service Manager partners closely with engineering, platform, infrastructure, security, and business teams to deliver stable, resilient, and high-performing services that align with business objectives and customer expectations. The role drives proactive risk management, operational maturity, automation, and service improvements while ensuring adherence to established operational processes and governance standards. The role also provides leadership across incident, problem, change, release, and service management disciplines for a team of Service Analysts to deliver resilient, secure, and customer-focused services that meet organizational goals.**Key responsibilities:*** Own the overall health, reliability, availability, and observability of critical business applications and technology services, ensuring alignment with established service commitments and customer expectations.* Establish, monitor, and report on Service Level Agreements (SLAs), Service Level Objectives (SLOs), Error Budgets, availability, performance, and operational KPIs to drive service excellence.* Lead Major Incident Management activities, coordinating cross-functional teams during outages, ensuring effective communication, rapid service restoration, and completion of Root Cause Analysis (RCA) for critical incidents.* Drive Problem Management practices by identifying recurring issues, analyzing systemic failures, implementing permanent corrective actions, and reducing operational risk through preventive measures.* Govern Change and Release Management processes by assessing operational risk, improving change success rates, supporting production readiness reviews, and coordinating maintenance and deployment activities.* Ensure effective observability across services through monitoring, alerting, logging, dashboards, and operational reporting that provide actionable insights into service health and performance.* Promote automation and operational efficiency by reducing manual processes, eliminating repetitive tasks, and advancing Infrastructure as Code (IaC), DevOps, self-healing, and auto-remediation capabilities.* Partner with engineering, platform, infrastructure, security, and business teams to proactively improve service resilience, scalability, stability, and customer experience.* Serve as the primary operational liaison for business stakeholders and vendors, providing regular service reviews, communicating risks, managing escalations, and ensuring alignment with business objectives.* Ensure adherence to ITIL-based operational processes, governance standards, compliance requirements, audit obligations, and the maintenance of operational documentation, runbooks, recovery procedures, and service support artifacts.* Lead continuous service improvement initiatives focused on reducing Mean Time to Recovery (MTTR), increasing service stability, improving customer satisfaction, and enhancing operational maturity.* Foster a culture of operational excellence, accountability, proactive risk management, customer focus, and continuous improvement across the service organization.* Own, lead, and facilitate operational governance ceremonies, including service reviews, incident and problem management reviews, change governance forums, operational readiness assessments, stakeholder communications, and executive service reporting.* Create, maintain, review, and ensure compliance with operational documentation, runbooks, standard operating procedures (SOPs), knowledge articles, audit evidence, disaster recovery procedures, and certification-related artifacts.* Provide leadership and oversight across incident, problem, change, release, and service management disciplines while ensuring consistent execution of operational processes and governance standards.* Lead, mentor, coach, and develop a team of Service Analysts, fostering technical growth, accountability, collaboration, and operational excellence.* Provide training for team members and users while assisting in building organizational bench strength through knowledge sharing, cross-training, and professional development initiatives.* Support operational risk management activities by identifying service vulnerabilities, assessing impacts, developing mitigation plans, and improving overall service resilience.* Collaborate with engineering and platform teams to improve service reliability, scalability, security, and production readiness while supporting modernization and transformation initiatives.* Drive strategic operational improvements that enhance service quality, customer experience, business alignment, and long-term operational sustainability.**Essential Skills:*** All applicants must demonstrate they have a legal right to work in the UK for employment at Allstate. Allstate is not providing sponsorship for this vacancy.* A minimum of 4 years of experience supporting, or improving enterprise technology services, infrastructure environments, platform operations, Site Reliability Engineering (SRE), IT Operations, or Service Management disciplines. (Or Equivalent).* A minimum of 2 years leading and mentoring teams within service reliability, availability, performance, and/or operational governance within a large enterprise environment.* Experience leading Major Incident response activities, coordinating cross-functional teams, and driving RCA efforts and corrective actions.* Experience implementing operational improvements, automation initiatives, risk-reduction measures, and continuous service improvement programs.* Strong understanding of observability practices, including monitoring, alerting, logging, dashboards, and operational reporting.* Knowledge of ITIL principles and IT Service Management (ITSM) processes.* Experience developing and maintaining operational documentation, runbooks, support procedures, recovery documentation, and knowledge articles.**Desirable Skills:*** Experience working within a SRE, DevOps, Cloud Operations, Platform Engineering, Enterprise Operations, or production support environment.* Experience supporting Identity and Access Management platforms, including IAM, ISAM, IBM Verify, SailPoint, or related identity technologies.* Experience managing Service Level Agreements (SLAs) and operational Key Performance Indicators (KPIs). (Or Equivalent)* Knowledge of networking technologies, firewalls, DNS, load balancing, and enterprise infrastructure concepts.* Experience leading or mentoring a team of engineers. (Or Equivalent)* Experience with Infrastructure as Code, automation frameworks, cloud-native operational practices, self-healing systems, or auto-remediation.* Familiarity with API management platforms and enterprise service-integration technologies.* Experience supporting messaging or event-streaming platforms such as Kafka.* Knowledge of middleware or integration technologies such as TIBCO or comparable enterprise platforms.* Experience supporting cloud platforms such as Microsoft Azure, Amazon Web Services, or Google Cloud Platform.* Experience using ServiceNow or a comparable ITSM platform.* Knowledge of compliance, risk management, audit controls, operational resilience, business continuity, and disaster recovery practices.* Relevant professional certifications, such as ITIL, SRE, cloud, Kubernetes, security, ServiceNow, or comparable technology certifications.* Experience leading operational maturity assessments, service governance programs, or production readiness reviews.**Job Posting End Date: Monday 21st September 2026 (11:59pm)**
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Service Manager – Site Reliability Engineering
Service Manager – Site Reliability Engineering

Jobtailor • Belfast City District

On-site
GBP 85,000 - 110,000
Managing Engineer, Data Security Engineering
Managing Engineer, Data Security Engineering

Allstate Northern Ireland Limited • Belfast City District

Hybrid
GBP 90,000 - 120,000
Managing Engineer - Observability, Pipeline & Analytics (Hybrid)
Managing Engineer - Observability, Pipeline & Analytics (Hybrid)

Allstate Insurance Company • Belfast City District

On-site
GBP 90,000 - 125,000
Corporate bonus scheme
Pension scheme
Annual performance-related pay reviews
+8
Managing Engineer. Cloud Security Engineering
Managing Engineer. Cloud Security Engineering

Allstate Insurance Company • Belfast City District

Hybrid
GBP 70,000 - 100,000
Corporate bonus scheme
Pension scheme
Annual performance-related pay reviews
+8
Managing Engineer - Infrastructure as Code (Hybrid)
Managing Engineer - Infrastructure as Code (Hybrid)

Allstate Insurance Company • Belfast City District

On-site
GBP 80,000 - 110,000
Hybrid working
Private medical and dental insurance
Two volunteering days
Product Security Engineers (Multiple Levels) Hybrid
Product Security Engineers (Multiple Levels) Hybrid

Allstate Northern Ireland • Belfast City District

Hybrid
GBP 70,000 - 110,000
Corporate bonus scheme
Pension scheme
Annual performance-related pay reviews
+7
Operations and SRE Manager
Operations and SRE Manager

LexisNexis Risk Solutions • Sutton

On-site
GBP 90,000 - 130,000
Senior Data Security Engineer - (Hybrid)
Senior Data Security Engineer - (Hybrid)

Allstate Insurance Company • Belfast City District

On-site
GBP 70,000 - 90,000
Corporate bonus scheme
Pension scheme
Annual performance reviews
+3
Observability SRE
Observability SRE

HCLTech • Greater London

On-site
GBP 70,000 - 95,000
Operations and SRE Manager
Operations and SRE Manager

LexisNexis Risk Solutions • United Kingdom

Remote
GBP 90,000 - 130,000