Lead Site Reliability Engineer: Observability & Resiliency

JPMorgan Chase & Co.

Glasgow

On-site

GBP 90,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. is seeking a Software Engineer to lead site reliability efforts in a world-class commercial and investment banking context. You will drive resiliency design reviews, mentor engineers, and break down complex problems for teams while guiding medium to large products.

Expect strong ownership and cross-team collaboration. You will champion reliability, improve service levels with data-driven analysis, and own incident response, architecture, and telemetry pipelines in hybrid

Qualifications

  • Formal training or certification on software engineering concepts.
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices.
  • Fluency in at least one programming language such as Java, Python, or Go.
  • Proficiency and experience in observability through monitoring, SLO alerting, and telemetry collection using industry tools.
  • Proficiency in CI/CD tools (e.g., Jenkins, GitLab, Terraform).
  • Experience with containers and container orchestration (ECS, Kubernetes, Docker).
  • Hands-on experience with OpenTelemetry collectors in production environments.
  • Ability to expand and collaborate across different levels and stakeholder groups.

Responsibilities

  • Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team
  • Leads initiatives to improve the reliability and stability of your team’s applications and platforms using data-driven analytics to improve service levels
  • Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers
  • Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology-related bottlenecks in your areas of expertise
  • Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses
  • Documents and shares knowledge within your organization via internal forums and communities of practice
  • Design, implement, and maintain operational reliability for large-scale Opentelemetry pipelines on hybrid on-prem/cloud environments.
  • Support telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheuse, Elasticsearch, OpenSearch optimizing for performance, real-time monitoring, logging, and alerting.
  • Assess, refactor, and incrementally migrate custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.

Skills

Site reliability
Observability
Java
Python
Go
CI/CD
Container orchestration
OpenTelemetry
Technical leadership

Education

Formal training in software engineering

Tools

Grafana
Prometheus
Datadog
Elasticsearch
OpenTelemetry

Job description

JPMorgan Chase & Co. is seeking a Software Engineer to lead site reliability efforts in a world-class commercial and investment banking context. You will drive resiliency design reviews, mentor engineers, and break down complex problems for teams while guiding medium to large products.

Expect strong ownership and cross-team collaboration. You will champion reliability, improve service levels with data-driven analysis, and own incident response, architecture, and telemetry pipelines in hybrid

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer - Observability & Resilience
Lead Site Reliability Engineer - Observability & Resilience

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 120,000
Lead Site Reliability Engineer – Observability & Resilience
Lead Site Reliability Engineer – Observability & Resilience

J.P. MORGAN • Scotland

Hybrid
GBP 90,000 - 130,000
Lead Site Reliability Engineer — Incident & Resilience
Lead Site Reliability Engineer — Incident & Resilience

J.P. MORGAN • Greater London

On-site
GBP 140,000 - 180,000
Lead Site Reliability Engineer - Resilience & Observability
Lead Site Reliability Engineer - Resilience & Observability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 150,000
Lead Site Reliability Engineer: Architect Resilience & AI Ops
Lead Site Reliability Engineer: Architect Resilience & AI Ops

JPMorgan Chase & Co. • City of Westminster

On-site
GBP 110,000 - 150,000
Senior SRE — Lead Resilience & Observability
Senior SRE — Lead Resilience & Observability

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer - Observability & Reliability
Lead Site Reliability Engineer - Observability & Reliability

JPMorganChase • Glasgow

On-site
GBP 70,000 - 90,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • City of Westminster

On-site
GBP 110,000 - 150,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000
Senior Site Reliability Engineer - AI-Driven Resilience Lead
Senior Site Reliability Engineer - AI-Driven Resilience Lead

J.P. MORGAN • Glasgow

On-site
GBP 110,000 - 160,000