Lead Site Reliability Engineer

JPMorgan Chase & Co.

Greater London

On-site

GBP 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. in London is seeking a Lead Site Reliability Engineer to shape next‑gen SRE patterns, observability, and reliability across globally distributed trading systems.

You will partner with front‑office traders, contribute production code (Java/Python/Kotlin), drive incident response, and lead AI‑assisted reliability initiatives while collaborating with infrastructure, cloud, and security teams.

Qualifications

  • Hands-on experience in front office trading or high‑pressure, low‑latency environments.
  • Proficiency with SRE tooling and techniques including FIX, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ, Oracle DB.
  • Experience using enterprise‑authorized AI capabilities to improve SRE workflows with validation and data sensitivity awareness.
  • Ability to evaluate AI recommendations for correctness and align with resiliency and security.
  • Deep knowledge of SLIs/SLOs, telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing observability frameworks for mission-critical systems.
  • Proven ability to lead incident response and drive long-term remediation.
  • Strong programming skills in Python, Java, or Kotlin; microservices and distributed/event-driven architectures.
  • Strong knowledge of CI/CD pipelines, automated testing, and deployment strategies.
  • Excellent communication with traders and senior stakeholders; calm under pressure.
  • Strong leadership presence with collaborative mindset.

Responsibilities

  • Engage daily with traders to understand workflows and reliability priorities.
  • Collaborate with desk to keep systems stable and performant.
  • Support live trading environments including incident response and post mortem leadership.
  • Contribute to codebase in Java, Kotlin, Python for reliability improvements.
  • Lead design and rollout of modern SRE patterns across trading systems.
  • Utilize AI capabilities to accelerate incident triage and analysis.
  • Champion AI-assisted reliability workflows with traceability and security controls.
  • Drive improvements in latency, throughput, and stability.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing.
  • Operate within a globally distributed engineering and trading organization.
  • Partner with infrastructure, cloud, cybersecurity teams to ensure end-to-end reliability.

Skills

SRE tooling
Front office trading
Python
Java
Kotlin
Observability
Incident response
CI/CD pipelines
Communication with traders
Leadership
Performance tuning

Tools

FIX messaging
Kafka
Grafana
Splunk
ITRS Geneos
Dynatrace
InfluxDB
MQ (IBM MQ)
Oracle DB

Job description

Our trading technology stack is undergoing a multi‑year convergence and modernization journey. You will play a pivotal role in shaping our next‑generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‑paced front‑office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization.

As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management’s Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front‑office trading platforms.

Job Responsibilities:
  • Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities.
  • Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs.
  • Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions.
  • Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation.
  • Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‑healing workflows, and resilience engineering.
  • Use enterprise‑authorized AI capabilities within the work environment to accelerate major‑incident triage, troubleshooting, and post‑incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Lead reuse‑first adoption of AI‑assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Drive improvements in latency, throughput, and stability across high volume trading applications.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments.
  • Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC.
  • Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability.
Required Qualifications, Capabilities, and Skills:
  • Strong hands‑on experience in front office trading environments or similarly high‑pressure, low‑latency domains.
  • Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB.
  • Demonstrated experience using enterprise‑authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI‑assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Deep knowledge of reliability engineering principles: SLIs/SLOs, real‑time telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing and implementing observability frameworks for mission critical systems.
  • Proven ability to lead incident response and drive long term remediation.
  • Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production‑grade code. Experience with microservices, distributed systems, and event‑driven architectures.
  • Strong understanding of CI/CD pipelines, automated testing, and deployment strategies.
  • Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations.
  • Strong leadership presence with a collaborative mindset.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 110,000 - 160,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Greater London

On-site
GBP 120,000 - 180,000
Lead Front-Office SRE for Trading Platforms
Lead Front-Office SRE for Trading Platforms

JPMorganChase • Greater London

On-site
GBP 120,000 - 180,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Lead Site Reliability / DevOps Engineer
Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 130,000
Front-Office SRE Lead: Observability & AI-Driven Reliability
Front-Office SRE Lead: Observability & AI-Driven Reliability

JPMorgan Chase & Co. • Greater London

On-site
GBP 120,000 - 160,000
Senior Lead SRE: Reliability, Observability & Resiliency
Senior Lead SRE: Reliability, Observability & Resiliency

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • City of Westminster

On-site
GBP 110,000 - 150,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000
Senior Lead Site Reliability / DevOps Engineer
Senior Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000