Lead SRE: AWS & Python for Scalable Reliability

JPMorganChase

Glasgow

On-site

GBP 90,000 - 120,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

JPMorganChase in the United Kingdom is seeking a Lead Site Reliability Engineer to drive reliability at scale across critical production systems. You will partner with software engineering and product teams to embed reliability into the software development lifecycle and deliver resilient services for millions of users.

You will lead incident response, define SLOs, champion observability, and mentor junior engineers, while shaping automation, CI/CD pipelines, and platform tooling to accelerate

Qualifications

  • Formal training or certification on site reliability engineering concepts and advanced applied experience.
  • Hands-on experience designing and operating large-scale distributed systems with a strong focus on availability, fault tolerance, and performance.
  • Proficiency in one or more programming or scripting languages (e.g., Python, Go, Java, Bash) for automation and tooling development.
  • Experience defining and managing service level indicators, service level objectives, and error budgets in production environments.
  • Strong background in observability tooling, including metrics, logging, and distributed tracing platforms.
  • Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements.
  • Experience with container orchestration and infrastructure-as-code practices (e.g., Kubernetes, Terraform, or equivalent).

Responsibilities

  • Lead the design and implementation of scalable, reliable, and observable infrastructure solutions that meet the firm's availability and performance standards.
  • Define and enforce service level objectives, error budgets, and reliability targets in partnership with engineering and product stakeholders.
  • Drive incident response, root cause analysis, and post-incident reviews to identify systemic improvements and reduce mean time to recovery.
  • Develop and maintain automation frameworks to eliminate toil, improve deployment pipelines, and accelerate delivery velocity.
  • Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process.
  • Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health.
  • Mentor and guide junior engineers, fostering a culture of continuous learning, operational discipline, and engineering excellence.
  • Evaluate and influence platform and tooling decisions to ensure alignment with reliability, scalability, and security requirements.
  • Leverage enterprise-authorized AI-assisted engineering practices to improve operational outcomes, including incident triage support, test strategy acceleration, and delivery workflow optimization, while ensuring consistent validation and secure handling of inputs and outputs.

Skills

Site reliability
Distributed systems
Programming skills
Observability
Incident management
Kubernetes

Tools

Kubernetes
Terraform

Job description

JPMorganChase in the United Kingdom is seeking a Lead Site Reliability Engineer to drive reliability at scale across critical production systems. You will partner with software engineering and product teams to embed reliability into the software development lifecycle and deliver resilient services for millions of users.

You will lead incident response, define SLOs, champion observability, and mentor junior engineers, while shaping automation, CI/CD pipelines, and platform tooling to accelerate

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead SRE: AWS & Python for Scalable Reliability
Lead SRE: AWS & Python for Scalable Reliability

JPMorgan Chase & Co. • United Kingdom

On-site
GBP 90,000 - 140,000
Lead SRE: AWS & Python for Scalable Reliability
Lead SRE: AWS & Python for Scalable Reliability

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE: AWS Platform & Reliability Leader
Lead SRE: AWS Platform & Reliability Leader

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Senior SRE: AWS Platform Reliability Architect
Senior SRE: AWS Platform Reliability Architect

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Senior SRE - AWS Platform & Reliability Lead
Senior SRE - AWS Platform & Reliability Lead

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer - Resilience & Observability
Lead Site Reliability Engineer - Resilience & Observability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 150,000
Senior Lead SRE - Reliability & Observability Leader
Senior Lead SRE - Reliability & Observability Leader

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer — SRE Leadership
Lead Site Reliability Engineer — SRE Leadership

Hackajob Ltd • Cumbernauld

Hybrid
GBP 85,000 - 130,000
Front-Office SRE Lead: Observability & AI-Driven Reliability
Front-Office SRE Lead: Observability & AI-Driven Reliability

JPMorgan Chase & Co. • Greater London

On-site
GBP 120,000 - 160,000