Lead Site Reliability Engineer

Arbuthnot Latham

England

On-site

GBP 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
BUPA Health cover
Discretionary bonus
Market-leading maternity/paternity policies

Job summary

A financial services company in the UK is seeking a Team Lead for Site Reliability Engineering to manage a NOC team transitioning to SRE. The role involves leading best practices, ensuring system reliability, and managing key performance indicators to enhance customer experience. Candidates should have a strong background in software engineering and team leadership with experience in automation and incident management tools.

Qualifications

  • Experience in leading root cause analysis and blameless postmortems.
  • Implementation of highly available and reliable systems.
  • Expertise with incident ticketing and change management systems.

Responsibilities

  • Lead, manage and motivate the NOC team.
  • Ensure the team follows best practices and audits quality.
  • Drive a culture of customer excellence and continual service improvement.

Skills

Line management/team leader experience
Understanding of software engineering principles
Expertise with monitoring and observability tools
Knowledge of cloud ecosystems
Knowledge of Infrastructure‑as‑code approaches
Knowledge of DevOps tools
SRE foundation course completed

Education

Bachelors or masters degree in computer science or related field

Tools

ServiceNow
AWS
Azure
Java
Python

Job description

Team Lead - Site Reliability Engineering

Join to apply for the Team Lead - Site Reliability Engineering role at Arbuthnot Latham

Arbuthnot Latham has been associated with banking since 1833. We combine private and commercial banking, wealth planning and investment management. We believe in traditional relationship and service-led banking powered by modern technology.

Job purpose

The Team Lead - Site Reliability Engineering is responsible for ensuring the effective and efficient running of the current NOC team with a view to transition to an SRE function over time. The team is responsible for enabling innovation and velocity of change while ensuring system reliability focusing on the critical features and functionality within products and platforms. It collaborates with the business or product owners to prioritise operational requirements by defining service‑level indicators (SLIs) and service‑level objectives (SLOs) to monitor and optimise customer journey and experience. Its goal is to design and operate scalable resilient systems utilising software engineering principles. It brings skills and expertise to automating manual tasks (TOIL) in such areas as incident management, problem management, change management, and release management tasks, and provides operational insights through monitoring and observability; and other aspects involved in preparing and optimising automated delivery solutions. To place the interests of customers at the centre of all activities, act in a way that is consistent with achieving good outcomes for consumers; and to comply with the FCA and PRA’s Conduct Rules.

Key Responsibilities
  • Lead, manage and motivate the team.
  • Ensure the team are following best practice across all disciplines.
  • Have oversight of team tasks including investigation, troubleshooting, diagnosis, resolution and recovery to minimise impact to services.
  • Audit the Engineers’ calls and tickets for quality assurance and provide feedback and coaching as required.
  • Drive a culture of Customer Excellence and Continual Service Improvement within the team.
  • Identify, develop, communicate, and implement process changes within the team.
  • Act as a point of escalation for the team.
SRE responsibilities
  • Help define the SRE practice for the organisation, collaborate with other stakeholders to select the relevant SRE principles, define the objectives and measurements of the outcomes.
  • Collaborate with stakeholders such as product and platform owners, to define service level objectives (SLOs), and service-level indicators (SLIs) for system operations focused on the critical features of the customers journey and experience.
  • Track and manage reliability performance against agreed SLOs, in partnership with other IT teams or other stakeholders, and ensure systems continue to meet SLOs over time.
  • Ensure key stakeholders, product owners, and platform owners are informed of reliability concerns and their potential impact to the customer experience.
  • Provide expert knowledge on reliability approaches, to ensure our organisation achieves its goals and roadmap for reliability.
  • Champion reliability being treated as a feature in products and platforms and promote the concept across all phases of the software development life cycle.
  • Create dashboards and reports to communicate key metrics, to product owners and key stakeholders.
  • Design, code, test and deliver solutions to automate manual operation (i.e., “TOIL”).
  • Participate in operations support and on‑call rotation shifts, for SRE supported systems and products.
  • Participate in or lead problem management activities, including post‑mortem incident analysis, and provision of technical insight, documented findings, outcomes and recommendations as part of a root cause analysis to troubleshoot priority incidents.
  • Implement automation to reduce probability and/or impact of problems recurring [possible options could include automated incident response, enhanced monitoring, observability initiatives, automation to change and release management].
  • Identify, evaluate, and recommend monitoring and observability tools and diagnostic techniques to improve system observability and insights, including identification of requirements, non‑functional requirements, design, implementation and operationalisation.
  • Participate in system design, platform management, capacity planning at launch reviews and sprint planning sessions, or product and platform architecture discussions. Ensure all operational requirements including availability, performance and disaster recovery are met.
  • Collaborate and share lessons learned regarding reliability, performance and incidents with all stakeholders.
  • Participate and exert influence in organisational learning initiatives such as communities of practice to share knowledge and foster a continuous learning and improvement mindset.
  • Support architects working on new solutions, including analysing requirements, supporting technical architecture activities, prototyping, designing and developing reusable infrastructure artifacts, testing, implementing, and preparing for ongoing support.
  • Train and mentor the team to ensure SRE best practices evolve and scale successfully in the organisation.
  • Shift working pattern - there is a requirement to work shifts and on‑call hours.
Knowledge/Experience/Skills
  • Line management/team leader experience
  • Understanding of software engineering principles (source control, versioning, code reviews, etc.)
  • Working in an environment that complies with ISO27001, NIST, CIS Benchmarks, PCIDSS amongst others
  • Leading root cause analysis and blameless postmortems in complex environments
  • Experience of communicating complex issues to senior stakeholders and technical teams.
  • Implementation of highly available and reliable systems, using multi‑AZ and multiregional approaches
  • Expertise with monitoring and observability tools (e.g. SolarWinds, Datadog, Azure/AWS native tools)
  • Expertise with SLI/SLO management tools such as (ServiceNow)
  • Expertise with Incident ticketing and change management systems such as (ServiceNow, Ivanti)
  • Expertise with automated incident response tools such as (Pager Duty, ServiceNow)
  • Expertise with software development frameworks/languages (e.g., Java, PHP, Python, PowerShell)
  • Extensive knowledge of cloud ecosystems (e.g. AWS, Red Hat OpenShift, Oracle Cloud Infrastructure, Microsoft Azure)
  • Knowledge of DevOps tools, such as CI/CD tools (e.g., Azure DevOps, GitHub, GitLab, Jira, Harness, Jenkins)
  • Knowledge of Infrastructure‑as‑code approaches, role‑specific automation tools and associated programming languages (e.g., AWS CloudFormation, Azure ARM, Hashi Corp Terraform, Progress Chef, Perforce Puppet)
  • Knowledge of Orchestration tools (e.g., Cloudify, env0, Morpheus Data, Pliant, RackN, Scalr, Spacelift, Terraform for Cloud) desirable
  • Cloud provider services (e.g., AWS, Azure, Oracle, regional providers)
  • Operating systems (e.g., Windows and Linux, including scripting experience)
  • Knowledge of scalable architectures, including APIs, microservices and PaaS desirable
  • Knowledge of architecting for resilience (e.g., HA, multi‑AZ, multiregional, backup and recovery tools) desirable
Qualifications
  • Bachelors or masters degree in computer science, information systems or a related field, or equivalent work experience
  • SRE foundation course completed, and qualification gained
  • Automation provider certifications
Life, Work and Benefits

Arbuthnot Latham is committed to equal‑opportunities for all staff and candidates. We embrace inclusion & diversity and understand why they are critical for the success of our business and people.

  • Competitive salary, pension & holiday allowance
  • BUPA Health cover
  • 4x Life Assurance
  • Discretionary bonus
  • Market‑leading maternity/paternity and menopause policies
Data Privacy and Reasonable adjustments

We take keeping your data security seriously. For more detail on how we may keep your data please refer to our Privacy Notice

https://careers.arbuthnotlatham.co.uk/files/RecruitmentPrivacyNotice.pdf

Reasonable adjustments: Please let us know of any adjustments or arrangements that you may need to help you apply to this role or that will help you during the recruitment process. If you wish to discuss any particular requirements or concerns you have because of a disability or medical condition please contact us at recruitment@arbuthnot.co.uk. Information you provide about any disability or medical condition will remain confidential unless it is necessary to disclose it to other members of staff or outside agencies to ensure the health and safety of yourself and others, or to implement the adjustments you require. In these circumstances we will first discuss with you how and to whom the information may be disclosed.

Seniority level
  • Mid‑Senior level
Employment type
  • Full‑time
Job function
  • Information Technology
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ServiceNow Application Specialist
ServiceNow Application Specialist

Arbuthnot Latham • England

On-site
GBP 50,000 - 70,000
Agile working environment
Competitive salary
Pension & holiday allowance
+3
Application Specialist - Banking Systems
Application Specialist - Banking Systems

Arbuthnot Latham • Greater London

Hybrid
GBP 55,000 - 75,000
One day WFH per week
Competitive holiday allowance
Pension
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

3761 Barclays - BX - UK • Knutsford

On-site
GBP 80,000 - 120,000
Head Of Infrastructure and Cloud - Internal Applicants Only
Head Of Infrastructure and Cloud - Internal Applicants Only

Arbuthnot Latham • Greater London

Hybrid
GBP 150,000 - 210,000
Competitive holiday allowance
Pension scheme
Discretionary bonus
+6
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 80,000 - 100,000
Healthcare
Retirement Planning
Paid Volunteering Days
+1
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Reward Gateway • Greater London

Hybrid
GBP 70,000 - 110,000
Life assurance
Pension
Employee Share Plan
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Barclays • Knutsford

On-site
GBP 90,000 - 140,000