Senior SRE - AI-Driven, Cloud & Observability Expert

Fairygodboss

Greater London

On-site

GBP 100,000 - 140,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

JPMorgan Chase is seeking a Site Reliability Engineer to join the International Consumer Bank team in the UK. You’ll focus on reliability, observability, and automation for customer-facing digital banking services, helping reduce operational toil and improve performance at scale.

You will work with Kubernetes, cloud services, and enterprise AI-assisted tooling to design for resilience, implement self-healing patterns, and develop actionable metrics and alerts across microservices.

Qualifications

  • Formal training or certification on software engineering concepts
  • Proven experience as a software engineer, including proficiency in at least one programming language such as Python, Go, or Java.
  • Demonstrated experience designing, coding, testing, and delivering software in at least one technology stack.
  • Strong debugging and troubleshooting skills across distributed systems.
  • Demonstrated experience as a Site Reliability Engineer or Site Reliability Engineer supporting production services.
  • Working knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing.
  • Experience with Kubernetes.
  • Experience with cloud computing services.
  • Familiarity with observability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger.
  • Ability to use AI-assisted engineering tools responsibly.
  • Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools.

Responsibilities

  • Drive continuous improvement of reliability, monitoring, and alerting for mission-critical microservices.
  • Reduce operational toil through automation by building reliable infrastructure and tooling that expedites feature development.
  • Develop meaningful service metrics, user journeys, service-level indicators, service-level objectives, error budgets, dashboards, and actionable alerts.
  • Engage with development teams throughout the software lifecycle to design for reliability and scale.
  • Design and implement self-healing and resiliency patterns, including graceful degradation, rate limiting, circuit breakers, and failover strategies.
  • Partner across engineering, product, and platform teams to promote reliability standards and adoption.
  • Execute performance testing and capacity planning to proactively identify and remove bottlenecks.
  • Participate in feature planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start.
  • Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation.
  • Continuously develop AI skills relevant to the role, including effective prompting, output validation, automation workflows, and safe usage patterns.
  • Drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes.
  • Applies knowledge of tools within the SDLC toolchain, including approved AI-assisted development and automation capabilities, to improve automation at scale.

Skills

Python
Go
Java
Distributed systems
Kubernetes
Cloud computing
AI-assisted tooling

Education

Bachelor's degree in Computer Science
Formal training in software engineering concepts

Tools

Grafana
Prometheus
Elasticsearch
Kibana
Jaeger

Job description

JPMorgan Chase is seeking a Site Reliability Engineer to join the International Consumer Bank team in the UK. You’ll focus on reliability, observability, and automation for customer-facing digital banking services, helping reduce operational toil and improve performance at scale.

You will work with Kubernetes, cloud services, and enterprise AI-assisted tooling to design for resilience, implement self-healing patterns, and develop actionable metrics and alerts across microservices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - AI-Driven Reliability
Senior Site Reliability Engineer - AI-Driven Reliability

Next Frontier Capital • Greater London

On-site
GBP 70,000 - 110,000
Senior Lead SRE - Reliability & Observability Leader
Senior Lead SRE - Reliability & Observability Leader

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Front-Office SRE Lead: Observability & AI-Driven Reliability
Front-Office SRE Lead: Observability & AI-Driven Reliability

JPMorgan Chase & Co. • Greater London

On-site
GBP 120,000 - 160,000
Senior SRE: AWS Platform Reliability Architect
Senior SRE: AWS Platform Reliability Architect

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Senior SRE — Lead Resilience & Observability
Senior SRE — Lead Resilience & Observability

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 140,000
Senior DevOps & SRE Lead AWS, CI/CD, Production Reliability
Senior DevOps & SRE Lead AWS, CI/CD, Production Reliability

JPMorganChase • Glasgow

On-site
GBP 120,000 - 160,000
Lead Site Reliability Engineer - Resilience & Observability
Lead Site Reliability Engineer - Resilience & Observability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 150,000
Lead SRE - Chase UK
Lead SRE - Chase UK

Next Frontier Capital • Greater London

On-site
GBP 70,000 - 110,000
Lead SRE - Chase UK
Lead SRE - Chase UK

Fairygodboss • Greater London

On-site
GBP 100,000 - 140,000
Senior SRE - AWS Platform & Reliability Lead
Senior SRE - AWS Platform & Reliability Lead

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000