- Own the overall health, reliability, availability, and observability of critical business applications and technology services
- Establish, monitor, and report on SLAs, SLOs, Error Budgets, availability, performance, and operational KPIs
- Lead Major Incident Management activities, coordinating cross-functional teams during outages and ensuring rapid service restoration and RCA completion
- Drive Problem Management by identifying recurring issues, analyzing systemic failures, implementing corrective actions, and reducing operational risk
- Govern Change and Release Management processes, production readiness reviews, maintenance, and deployment activities
- Ensure effective observability through monitoring, alerting, logging, dashboards, and operational reporting
- Promote automation and operational efficiency through Infrastructure as Code, DevOps, self-healing, and auto-remediation
- Partner with engineering, platform, infrastructure, security, and business teams to improve resilience, scalability, stability, and customer experience
- Act as operational liaison for business stakeholders and vendors, conducting service reviews, communicating risks, and managing escalations
- Ensure adherence to ITIL-based processes, governance standards, compliance requirements, audit obligations, and operational documentation standards
- Lead continuous service improvement initiatives to reduce MTTR, increase stability, improve customer satisfaction, and enhance operational maturity
- Lead operational governance ceremonies, service reviews, incident and problem reviews, change governance forums, readiness assessments, stakeholder communications, and executive service reporting
- Create, maintain, review, and ensure compliance with runbooks, SOPs, knowledge articles, audit evidence, disaster recovery procedures, and certification artifacts
- Provide leadership across incident, problem, change, release, and service management disciplines
- Lead, mentor, coach, and develop a team of Service Analysts
- Provide training, knowledge sharing, cross-training, and professional development support
- Support operational risk management by identifying vulnerabilities, assessing impacts, and developing mitigation plans
- Collaborate on service reliability, scalability, security, production readiness, modernization, and transformation initiatives
- Drive strategic operational improvements supporting service quality, customer experience, business alignment, and long-term sustainability
Requirements
- Legal right to work in the UK; Allstate is not providing sponsorship for this vacancy
- Minimum of 4 years of experience supporting or improving enterprise technology services, infrastructure environments, platform operations, Site Reliability Engineering, IT Operations, or Service Management disciplines (or equivalent)
- Minimum of 2 years leading and mentoring teams within service reliability, availability, performance, and/or operational governance in a large enterprise environment
- Experience leading Major Incident response activities, coordinating cross-functional teams, and driving RCA efforts and corrective actions
- Experience implementing operational improvements, automation initiatives, risk-reduction measures, and continuous service improvement programs
- Strong understanding of observability practices, including monitoring, alerting, logging, dashboards, and operational reporting
- Knowledge of ITIL principles and IT Service Management processes
- Experience developing and maintaining operational documentation, runbooks, support procedures, recovery documentation, and knowledge articles
- Experience working within an SRE, DevOps, Cloud Operations, Platform Engineering, Enterprise Operations, or production support environment
- Experience supporting Identity and Access Management platforms, including IAM, ISAM, IBM Verify, SailPoint, or related identity technologies
- Experience managing SLAs and operational KPIs (or equivalent)
- Knowledge of networking technologies, firewalls, DNS, load balancing, and enterprise infrastructure concepts
- Experience leading or mentoring a team of engineers (or equivalent)
- Experience with Infrastructure as Code, automation frameworks, cloud-native operational practices, self-healing systems, or auto-remediation
- Familiarity with API management platforms and enterprise service-integration technologies
- Experience supporting messaging or event-streaming platforms such as Kafka
- Knowledge of middleware or integration technologies such as TIBCO or comparable enterprise platforms
- Experience supporting Microsoft Azure, Amazon Web Services, or Google Cloud Platform
- Experience using ServiceNow or a comparable ITSM platform
- Knowledge of compliance, risk management, audit controls, operational resilience, business continuity, and disaster recovery practices
- Relevant professional certifications, such as ITIL, SRE, cloud, Kubernetes, security, ServiceNow, or comparable technology certifications
- Experience leading operational maturity assessments, service governance programs, or production readiness reviews
Core Competencies
Demonstrates expertise in Service Reliability Engineering, Incident Management, and ITIL-based processes, with a strong focus on operational governance, automation, and continuous service improvement. Proven ability to lead cross-functional teams, manage SLAs, and enhance customer experience through effective communication and collaboration.
Highest-signal resume keywords
- Service Reliability Engineering
- Incident Management
- ITIL Principles
- Infrastructure as Code
- Operational Governance
Hard Skills
- Operational Documentation
- Monitoring
- Alerting
- Logging
- Dashboards
- Cloud Operations
- Automation Frameworks
- Identity and Access Management
- Networking Technologies
- Event-Streaming Platforms
Soft Skills
- Leadership
- Mentoring
- Collaboration
- Communication
- Problem-Solving
Certifications & Qualifications
- ITIL
- SRE
- Cloud Certifications
- Kubernetes
- Security Certifications
Industry Keywords
- Operational Resilience
- Business Continuity
- Disaster Recovery
- Compliance
- Risk Management
Tools & Technologies
- ServiceNow
- Microsoft Azure
- Amazon Web Services
- Google Cloud Platform
- Kafka
- TIBCO