Lead Site Reliability Engineer - Resilience & Observability

JPMorganChase

Glasgow

On-site

GBP 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase in the United Kingdom seeks a Lead Site Reliability Engineer to shape reliability for large-scale services. You will lead the team, drive resiliency design reviews, and mentor engineers while guiding open telemetry pipelines across hybrid environments.

The role focuses on incident leadership, SLOs, observability with Grafana/Prometheus, and continuous improvement of production systems. You will collaborate with stakeholders to ensure stable, scalable platforms and minimize outages.

Qualifications

  • Formal training or certification on software engineering concepts and advanced applied experience.
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform.
  • Fluency in at least one programming language such as Java, Python, Go, etc.
  • Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc.
  • Proficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, Terraform, etc.
  • Experience with container and container orchestration such as ECS, Kubernetes, Docker, etc.
  • Hands‑on experience with the design, deployment, and operation of OpenTelemetry collectors in production environments, focusing on technical aspects such as configuring, optimizing, and troubleshooting OTLP endpoints and receivers.
  • Ability to expand and collaborate across different levels and stakeholder groups

Responsibilities

  • Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team
  • Leads initiatives to improve the reliability and stability of your team's applications and platforms using data‑driven analytics to improve service levels
  • Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers
  • Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology‑related bottlenecks in your areas of expertise
  • Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses
  • Documents and shares knowledge within your organization via internal forums and communities of practice
  • Design, implement, and maintain operational reliability for large‑scale OpenTelemetry pipelines on hybrid on‑prem/cloud environments.
  • Support telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, OpenSearch optimizing for performance, real‑time monitoring, logging, and alerting.
  • Assess, refactor, and incrementally migrate custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.

Skills

Java
Python
Go
Observability
Kubernetes
Docker
OpenTelemetry
SRE practices

Education

Software engineering certification

Tools

Jenkins
GitLab
Terraform
Prometheus
Grafana
ELK

Job description

JPMorgan Chase in the United Kingdom seeks a Lead Site Reliability Engineer to shape reliability for large-scale services. You will lead the team, drive resiliency design reviews, and mentor engineers while guiding open telemetry pipelines across hybrid environments.

The role focuses on incident leadership, SLOs, observability with Grafana/Prometheus, and continuous improvement of production systems. You will collaborate with stakeholders to ensure stable, scalable platforms and minimize outages.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer - Observability & Resilience
Lead Site Reliability Engineer - Observability & Resilience

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 120,000
Senior Lead SRE - Reliability & Observability Leader
Senior Lead SRE - Reliability & Observability Leader

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Lead Site Reliability Engineer: Observability & Resiliency
Lead Site Reliability Engineer: Observability & Resiliency

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer - Observability & Reliability
Lead Site Reliability Engineer - Observability & Reliability

JPMorganChase • Glasgow

On-site
GBP 70,000 - 90,000
Lead Site Reliability Engineer: Architect Resilience & AI Ops
Lead Site Reliability Engineer: Architect Resilience & AI Ops

JPMorgan Chase & Co. • City of Westminster

On-site
GBP 110,000 - 150,000
Lead Site Reliability Engineer – Observability & Resilience
Lead Site Reliability Engineer – Observability & Resilience

J.P. MORGAN • Scotland

Hybrid
GBP 90,000 - 130,000
Senior SRE — Lead Resilience & Observability
Senior SRE — Lead Resilience & Observability

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 140,000
Lead SRE: AWS Platform & Reliability Leader
Lead SRE: AWS Platform & Reliability Leader

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Lead Site Reliability Engineer — Incident & Resilience
Lead Site Reliability Engineer — Incident & Resilience

J.P. MORGAN • Greater London

On-site
GBP 140,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000