Lead Site Reliability Engineer - Glasgow

JP Morgan Chase

Glasgow

On-site

GBP 62,000 - 102,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JP Morgan Chase is seeking a Senior Reliability Engineer in the United Kingdom to lead incident management, drive resilience across platforms, and partner with cross-functional teams to improve operational outcomes.

The role focuses on enterprise reliability programs, governance, and the use of AI-assisted tooling to提升 code quality, delivery speed, and security within production environments. Strong ownership and urgency are required.

Qualifications

  • Formal training or certification on software engineering concepts.
  • Experience improving operational processes, including incident management, root cause analysis, problem management, and governance.
  • Strong knowledge of reliability engineering concepts including SLOs, indicators, error budgets, capacity planning, resilience, and observability.
  • Experience with cloud platforms and infrastructure-as-code tooling such as Terraform.
  • Proven ability to lead reliability programs across multiple services or teams.

Responsibilities

  • Own and continuously improve the incident management lifecycle with triage, escalation, communications, and recovery.
  • Serve as Incident Commander for major incidents with clear stakeholder updates.
  • Drive drills, game days, and failure-mode exercises to improve readiness and resilience.
  • Define, implement, and enforce reliability standards, including SLOs, monitoring, and alert quality.
  • Strengthen change and release management with readiness checks and rollback procedures.
  • Partner with Engineering, Product, Infrastructure, and Security to prioritise reliability work.
  • Lead problem management and root cause analysis to identify systemic fixes.
  • Own reliability reporting and governance with KPI tracking and executive summaries.
  • Facilitate operations reviews, incident boards, and reliability councils.
  • Promote enterprise AI-assisted engineering practices with validated outputs and safe usage.
  • Apply SDLC tools to improve value from automation and tooling.

Skills

Incident management
Root cause analysis
Problem management
Operational governance
Service level objectives
Observability
Cross-functional collaboration
Analytics & reporting
Cloud platforms
AI-assisted engineering tools
Automation & runbook design

Education

Certification in software engineering concepts

Tools

Terraform

Job description

Salary: £62,000 - 102,000 per year

Requirements
  • Formal training or certification on software engineering concepts and advanced applied experience.
  • Demonstrated experience improving operational processes, including incident management, root cause analysis, problem management, and operational governance in complex production environments.
  • Strong knowledge of reliability engineering concepts including service level objectives and indicators, error budgets, capacity planning, resilience patterns, and observability.
  • Proven ability to influence across teams, drive cross-functional initiatives, and communicate clearly with stakeholders, particularly in high-urgency, time-sensitive situations.
  • Strong analytical and reporting skills with the ability to define metrics, generate actionable insights, and drive decisions based on operational data.
  • Hands-on experience with cloud platforms such as Amazon Web Services, Google Cloud Platform, or Microsoft Azure, and infrastructure-as-code tooling including Terraform.
  • Systematic problem-solving and troubleshooting skills applied to complex, distributed systems in production environments.
  • Demonstrated ability to operate with a high degree of ownership, self-direction, and urgency in ambiguous or fast-moving situations.
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment, with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs and outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.
  • Experience leading reliability programs across multiple services or teams at a platform or domain level.
  • Experience building automation for operations, including runbook automation, event correlation, alert tuning, and self-healing patterns.
  • Exposure to agentic AI or AI-driven operations concepts, including workflow orchestration, decision support, and safe automation patterns.
  • Familiarity with operational governance frameworks and experience facilitating reliability forums or review boards at scale.
Responsibilities
  • Own and continuously improve the incident management lifecycle, including triage, escalation, stakeholder communications, and recovery, ensuring consistent execution and measurable improvement over time.
  • Serve as Incident Commander for major incidents, coaching responders to follow defined processes and site reliability engineering best practices while maintaining clear and timely stakeholder communication.
  • Drive operational readiness through drills, game days, and failure-mode exercises that improve response effectiveness, surface gaps, and build team resilience before incidents occur.
  • Define, implement, and enforce reliability standards across services, including service level indicators and objectives, error budgets, monitoring coverage, and alert quality.
  • Strengthen change and release management practices by establishing readiness checks, progressive delivery standards, rollback procedures, runbook quality, and production hygiene expectations.
  • Partner with Engineering, Product, Infrastructure, and Security teams to prioritize reliability work and deliver solutions that address user and operational pain points with lasting impact.
  • Lead problem management and root cause analysis by debugging complex production issues, identifying systemic root causes, and driving durable remediation and preventative actions.
  • Own reliability reporting and operational governance by tracking key performance indicators, including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate, and producing executive-ready summaries.
  • Facilitate operational forums including operations reviews, incident review boards, and reliability councils to drive accountability, share learnings, and align stakeholders on priorities.
  • Drive team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes, while establishing consistent validation standards and promoting reuse of effective patterns across the team.
  • Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Technologies
  • Agentic AI
  • AI
  • Azure
  • Cloud
  • Incident Management
  • Support
  • Marketing
  • Security
  • Terraform
  • Web
More

We are JPMorganChase, a global leader in financial services, providing strategic advice and products to corporations, governments, wealthy individuals, and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do, and we strive to build trusted, long-term partnerships that help our clients achieve their business objectives. We value the diverse talents of our people and are committed to diversity and inclusion, with equal opportunity practices and reasonable accommodations for applicants and employees. Our Corporate Functions teams span finance, risk, human resources, and marketing, helping set our businesses, clients, customers, and employees up for success. This is a full-time role in our AI/ML & Data Platforms area, and the posting date is 2026-07-23.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE - AWS Platform
Lead SRE - AWS Platform

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000
Lead Site Reliability Engineer – Operations Excellence
Lead Site Reliability Engineer – Operations Excellence

J.P. MORGAN • Bournemouth

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer – Operations Excellence
Lead Site Reliability Engineer – Operations Excellence

J.P. MORGAN • Greater London

On-site
GBP 140,000 - 180,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer – Operations Excellence
Lead Site Reliability Engineer – Operations Excellence

J.P. MORGAN • Glasgow

On-site
GBP 110,000 - 160,000
Software Engineer III - AI/ML Platform Reliability
Software Engineer III - AI/ML Platform Reliability

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • City of Westminster

On-site
GBP 110,000 - 150,000
Software Engineer III - AI/ML Platform Reliability
Software Engineer III - AI/ML Platform Reliability

JPMorganChase • Glasgow

On-site
GBP 85,000 - 120,000
Lead Software Engineer - Glasgow
Lead Software Engineer - Glasgow

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer - Operations Excellence
Lead Site Reliability Engineer - Operations Excellence

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000