Lead SRE: AWS & Python for Scalable Reliability

Next Frontier Capital

Glasgow

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

JPMorganChase is seeking a Lead Site Reliability Engineer to ensure the availability, performance, and resilience of production systems serving millions globally. You will embed reliability in the software lifecycle, guide cross-functional teams, and champion observability across monitoring and incident practices.

The role emphasizes leadership in incident response, SLOs, and tooling, with a focus on engineering excellence and scalable, secure operations in a financial services environment.

Qualifications

  • Formal training or certification on site reliability engineering concepts and advanced applied experience.
  • Hands-on experience designing and operating large-scale distributed systems with a strong focus on availability, fault tolerance, and performance.
  • Proficiency in programming or scripting languages (Python, Go, Java, Bash) for automation and tooling development.
  • Experience defining and managing service level indicators, service level objectives, and error budgets in production environments.
  • Strong background in observability tooling, including metrics, logging, and distributed tracing platforms.
  • Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements.
  • Experience with container orchestration and infrastructure-as-code practices (Kubernetes, Terraform, or equivalent).
  • Ability to communicate complex technical concepts clearly to both technical and non-technical stakeholders.

Responsibilities

  • Lead the design and implementation of scalable, reliable, and observable infrastructure solutions that meet the firm's availability and performance standards.
  • Define and enforce service level objectives, error budgets, and reliability targets in partnership with engineering and product stakeholders.
  • Drive incident response, root cause analysis, and post-incident reviews to identify systemic improvements and reduce mean time to recovery.
  • Develop and maintain automation frameworks to eliminate toil, improve deployment pipelines, and accelerate delivery velocity.
  • Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process.
  • Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health.
  • Mentor and guide junior engineers, fostering a culture of continuous learning, operational discipline, and engineering excellence.
  • Evaluate and influence platform and tooling decisions to ensure alignment with reliability, scalability, and security requirements.
  • Leverage enterprise-authorized AI-assisted engineering practices to improve operational outcomes, including incident triage support, test strategy acceleration, and delivery workflow optimization, while ensuring consistent validation and secure handling of inputs and outputs.

Skills

Python
Go
Java
Automation
Incident response
Observability
Communication

Education

SRE concepts certification

Tools

Kubernetes
Terraform
Prometheus
Grafana

Job description

JPMorganChase is seeking a Lead Site Reliability Engineer to ensure the availability, performance, and resilience of production systems serving millions globally. You will embed reliability in the software lifecycle, guide cross-functional teams, and champion observability across monitoring and incident practices.

The role emphasizes leadership in incident response, SLOs, and tooling, with a focus on engineering excellence and scalable, secure operations in a financial services environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead SRE: AWS & Python for Scalable Reliability
Lead SRE: AWS & Python for Scalable Reliability

JPMorgan Chase & Co. • United Kingdom

On-site
GBP 90,000 - 140,000
Lead SRE: AWS & Python for Scalable Reliability
Lead SRE: AWS & Python for Scalable Reliability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 120,000
Lead SRE: AWS Platform & Reliability Leader
Lead SRE: AWS Platform & Reliability Leader

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Senior SRE: AWS Platform Reliability Architect
Senior SRE: AWS Platform Reliability Architect

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Senior SRE - AWS Platform & Reliability Lead
Senior SRE - AWS Platform & Reliability Lead

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer — SRE Leadership
Lead Site Reliability Engineer — SRE Leadership

Hackajob Ltd • Cumbernauld

Hybrid
GBP 85,000 - 130,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer: Architect Resilient Systems
Lead Site Reliability Engineer: Architect Resilient Systems

Hackajob Ltd • Milton

On-site
GBP 110,000 - 140,000
Front-Office SRE Lead: Observability & AI-Driven Reliability
Front-Office SRE Lead: Observability & AI-Driven Reliability

JPMorgan Chase & Co. • Greater London

On-site
GBP 120,000 - 160,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000