Lead SRE - AWS,Python

JPMorgan Chase & Co.

Glasgow

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

JPMorgan Chase & Co. is seeking a Lead Site Reliability Engineer to ensure availability, performance, and resilience of production systems serving millions of customers globally.

You will partner with engineering and product teams to embed reliability practices in the software lifecycle, driving a culture of operational excellence and scalable, secure infrastructure.

Qualifications

  • Formal training or certification in site reliability engineering concepts.
  • Experience operating large-scale distributed systems with emphasis on availability and performance.
  • Programming or scripting in Python/Go/Java/Bash for automation.
  • Experience defining and managing SLIs/SLOs and error budgets in production.
  • Strong observability background with metrics, logs, and distributed tracing.

Responsibilities

  • Lead design and implementation of scalable, reliable infrastructure.
  • Define and enforce SLIs, SLOs, and error budgets with stakeholders.
  • Drive incident response, blameless post-mortems, and systemic improvements.
  • Develop automation frameworks to reduce toil and accelerate delivery.
  • Collaborate with software, architecture, and security teams on reliability.

Skills

Site Reliability Engineering
Distributed Systems
Observability
Incident Response
Kubernetes
Terraform
Automation
Python/Go/Java

Tools

Prometheus
ELK/Logging
Tracing (OpenTelemetry)

Job description

Drive reliability at scale — join a team where your engineering expertise shapes the resilience of critical systems.

JPMorganChase is one of the world’s leading financial services firms, and the technology that powers it demands the highest standards of reliability, performance, and scale. Here, you will work alongside talented engineers who are passionate about building systems that never sleep — and you will have the opportunity to grow your career while solving some of the most complex infrastructure challenges in the industry. We invest in our people, our platforms, and our future — and we want you to be part of it.

As a Lead Site Reliability Engineer at JPMorganChase, you will play a critical role in ensuring the availability, performance, and resilience of production systems that serve millions of customers and clients globally. You will partner with engineering and product teams to embed reliability practices into the software development lifecycle, driving a culture of operational excellence. Your work will directly impact the firm’s ability to deliver seamless, uninterrupted services at enterprise scale.

Job responsibilities
  • Lead the design and implementation of scalable, reliable, and observable infrastructure solutions that meet the firm’s availability and performance standards
  • Define and enforce service level objectives, error budgets, and reliability targets in partnership with engineering and product stakeholders
  • Drive incident response, root cause analysis, and post-incident reviews to identify systemic improvements and reduce mean time to recovery
  • Develop and maintain automation frameworks to eliminate toil, improve deployment pipelines, and accelerate delivery velocity
  • Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process
  • Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health
  • Mentor and guide junior engineers, fostering a culture of continuous learning, operational discipline, and engineering excellence
  • Evaluate and influence platform and tooling decisions to ensure alignment with reliability, scalability, and security requirements
  • Leverage enterprise-authorized AI-assisted engineering practices to improve operational outcomes, including incident triage support, test strategy acceleration, and delivery workflow optimization, while ensuring consistent validation and secure handling of inputs and outputs
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and advanced applied experience
  • Hands-on experience designing and operating large-scale distributed systems with a strong focus on availability, fault tolerance, and performance
  • Proficiency in one or more programming or scripting languages (e.g., Python, Go, Java, Bash) for automation and tooling development
  • Experience defining and managing service level indicators, service level objectives, and error budgets in production environments
  • Strong background in observability tooling, including metrics, logging, and distributed tracing platforms
  • Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements
  • Experience with container orchestration and infrastructure-as-code practices (e.g., Kubernetes, Terraform, or equivalent)
  • Ability to communicate complex technical concepts clearly to both technical and non-technical stakeholders
Preferred qualifications, capabilities, and skills
  • Experience operating in a regulated financial services or similarly complex enterprise environment
  • Familiarity with chaos engineering principles and tools used to proactively test system resilience
  • Exposure to cloud-native architectures and multi-cloud or hybrid infrastructure environments
  • Experience contributing to platform engineering or internal developer tooling initiatives
  • Background in capacity planning, performance engineering, or cost optimization at scale
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE - AWS,Python
Lead SRE - AWS,Python

JPMorganChase • Glasgow

On-site
GBP 90,000 - 120,000
Lead Software Engineer – Software Reliability
Lead Software Engineer – Software Reliability

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Lead Software Engineer - Software Reliability
Lead Software Engineer - Software Reliability

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 120,000
Lead SRE - AWS,Python
Lead SRE - AWS,Python

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Senior Lead Site Reliability / DevOps Engineer
Senior Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000