Lead Site Reliability Engineer - Observability & Resilience

JPMorgan Chase & Co.

Glasgow

On-site

GBP 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. seeks a Lead Site Reliability Engineer to define the future of reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers across large-scale OpenTelemetry pipelines in hybrid environments.

You will guide incident responses, drive data‑driven improvements to SLOs and error budgets, and influence cross‑team decisions with a strong emphasis on performance, security, and scalable architecture.

Qualifications

  • Formal training or certification on software engineering concepts and advanced applied experience.
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform.
  • Fluency in at least one programming language such as Java, Python, Go, etc.
  • Proficiency and experience in observability with white/black box monitoring and SLO alerting using Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc.
  • Proficiency in CI/CD tools (Jenkins, GitLab, Terraform).
  • Experience with container and orchestration (ECS, Kubernetes, Docker).
  • Hands-on with OpenTelemetry collectors in production; configuring OTLP endpoints and receivers.
  • Ability to expand and collaborate across different levels and stakeholder groups.

Responsibilities

  • Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team
  • Leads initiatives to improve the reliability and stability of your team’s applications and platforms using data‑driven analytics to improve service levels
  • Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers
  • Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology‑related bottlenecks in your areas of expertise
  • Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses
  • Documents and shares knowledge within your organization via internal forums and communities of practice
  • Design, implement, and maintain operational reliability for large‑scale Opentelemetry pipelines on hybrid on‑prem/cloud environments.
  • Support telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, OpenSearch optimizing for performance, real‑time monitoring, logging, and alerting.
  • Assess, refactor, and incrementally migrate custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.

Skills

Reliability
Scalability
Performance
Security
Enterprise architecture
Toil reduction
SRE best practices
Java
Python
Go
Grafana
Dynatrace
Prometheus
Datadog
Splunk
Elasticsearch
Jenkins
GitLab
Terraform
Docker
Kubernetes
ECS
OpenTelemetry
OpAMP
Stakeholder collaboration
Distributed tracing
Metrics
Logging
AWS certification
Kubernetes certification
Open-source contributions

Education

Formal training or certification in software engineering concepts

Tools

OpenTelemetry

Job description

JPMorgan Chase & Co. seeks a Lead Site Reliability Engineer to define the future of reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers across large-scale OpenTelemetry pipelines in hybrid environments.

You will guide incident responses, drive data‑driven improvements to SLOs and error budgets, and influence cross‑team decisions with a strong emphasis on performance, security, and scalable architecture.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer: Observability & Resiliency
Lead Site Reliability Engineer: Observability & Resiliency

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer - Resilience & Observability
Lead Site Reliability Engineer - Resilience & Observability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 150,000
Lead Site Reliability Engineer – Observability & Resilience
Lead Site Reliability Engineer – Observability & Resilience

J.P. MORGAN • Scotland

Hybrid
GBP 90,000 - 130,000
Senior SRE — Lead Resilience & Observability
Senior SRE — Lead Resilience & Observability

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer: Architect Resilience & AI Ops
Lead Site Reliability Engineer: Architect Resilience & AI Ops

JPMorgan Chase & Co. • City of Westminster

On-site
GBP 110,000 - 150,000
Lead Site Reliability Engineer — Incident & Resilience
Lead Site Reliability Engineer — Incident & Resilience

J.P. MORGAN • Greater London

On-site
GBP 140,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000
Senior Lead SRE - Reliability & Observability Leader
Senior Lead SRE - Reliability & Observability Leader

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Lead Site Reliability / DevOps Engineer
Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer - Observability & Reliability
Lead Site Reliability Engineer - Observability & Reliability

JPMorganChase • Glasgow

On-site
GBP 70,000 - 90,000