Senior Lead Site Reliability / DevOps Engineer

JP Morgan Chase

Glasgow

On-site

GBP 62,000 - 102,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

J.P. Morgan Chase is seeking a Senior Site Reliability/Observability Engineer to lead reliability design and implement large-scale OpenTelemetry pipelines in hybrid environments.

You will drive incident response, define SLIs/SLIs, and push observability improvements across teams in a global financial services setting. The role requires deep cloud-native experience, advanced expertise in monitoring and tracing, and collaboration with multiple stakeholder groups to ensure high-availability

Qualifications

  • Formal training or certification in software engineering concepts and advanced experience in system design, application development, testing, and operational stability.
  • Advanced knowledge of reliability, scalability, performance, security, enterprise architecture, toil reduction, and site reliability best practices.
  • Advanced proficiency in Java, Python, or Go.
  • Advanced proficiency with observability, SLO alerting, and telemetry collection using Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch.
  • Proficiency with CI/CD tools such as Jenkins, GitLab, Terraform.
  • Experience with containers and orchestration (ECS, Kubernetes, Docker).
  • Hands-on experience with OpenTelemetry collectors in production, OTLP endpoints.
  • Ability to solve reliability design problems independently.
  • Practical cloud-native experience.
  • Ability to collaborate with stakeholders.

Responsibilities

  • Provide technical guidance and direction on site reliability practices to support business, technical teams, contractors, and vendors.
  • Develop secure, high-quality production code for reliability tooling and telemetry pipelines, and review code.
  • Drive decisions influencing reliability design, observability architecture, and operational processes.
  • Serve as a subject matter expert in site reliability, observability, or telemetry engineering.
  • Lead resiliency design reviews and break complex problems into manageable work for other engineers.
  • Act as main contact during major incidents, resolve issues to avoid losses, and promote blameless postmortems.
  • Collaborate to define service level indicators, objectives, and error budgets.
  • Design and maintain operational reliability for OpenTelemetry pipelines in hybrid environments, exporting to backends like InfluxDB, Prometheus, Elasticsearch, OpenSearch.
  • Drive migration of legacy telemetry to OpenTelemetry instrumentation while preserving stability.
  • Contribute to engineering community and advocate best practices for observability and reliability.

Skills

Software engineering concepts
System design
Application development
Testing
Operational stability
Reliability engineering
Scalability
Performance tuning
Security
Enterprise architecture
Toil reduction
Cloud
Observability
Distributed systems
Java
Python
Go
Monitoring
SRE tooling
SLOs
Telemetry collection
OpenTelemetry
OpenTelemetry instrumentation

Tools

AWS
Kubernetes
Docker
ECS
GitLab
Terraform
OpenTelemetry
Prometheus
Datadog
Grafana
Dynatrace
Elasticsearch
OpenSearch
Jenkins

Job description

Salary: £62,000 - 102,000 per year

Requirements
  • We require formal training or certification in software engineering concepts, along with advanced applied experience delivering system design, application development, testing, and operational stability.
  • We require advanced knowledge of reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with in-depth expertise in one or more technical disciplines such as cloud, observability, or distributed systems.
  • We require advanced proficiency in one or more programming languages such as Java, Python, or Go.
  • We require advanced proficiency and experience in observability, including white-box and black-box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, and similar platforms.
  • We require proficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, and Terraform.
  • We require experience with containers and container orchestration technologies such as ECS, Kubernetes, and Docker.
  • We require hands‑on experience designing, deploying, and operating OpenTelemetry collectors in production, including configuring, optimizing, and troubleshooting OTLP endpoints and receivers.
  • We require the ability to solve reliability design and functionality problems independently with little to no oversight.
  • We require practical cloud‑native experience.
  • We require the ability to collaborate effectively across different levels and stakeholder groups.
  • Preferred: knowledge of distributed tracing, metrics, and logging best practices.
  • Preferred: certification in AWS, Kubernetes, or relevant technologies.
  • Preferred: proven track record in system health monitoring, capacity management, and blameless postmortems for high‑availability services.
  • Preferred: deep understanding of distributed system design principles, networking concepts such as TCP/IP, DNS, and load balancing, and Linux internals.
  • Preferred: contributions to open‑source observability or telemetry projects.
  • Preferred: experience with agent control planes and management protocols, with hands‑on knowledge of OpAMP highly desirable.
Responsibilities
  • We provide technical guidance and direction on site reliability practices to support our business, technical teams, contractors, and vendors.
  • We develop secure, high-quality production code for reliability tooling and telemetry pipelines, and review and debug code written by others.
  • We drive decisions that influence reliability design, observability architecture, application functionality, and technical operations and processes.
  • We serve as a subject matter expert in one or more areas of site reliability, observability, or telemetry engineering.
  • We lead resiliency design reviews and break complex reliability problems into digestible work for other engineers, acting as a technical lead for large products.
  • We act as the main point of contact during major incidents, identify and solve issues quickly to avoid financial losses, and champion a blameless postmortem culture.
  • We collaborate with team members and stakeholders to define service level indicators, service level objectives, and error budgets.
  • We design, implement, and maintain operational reliability for large‑scale OpenTelemetry pipelines in hybrid on‑prem and cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch.
  • We drive the assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.
  • We actively contribute to the engineering community as advocates of firmwide frameworks, tools, and practices, and influence peers and project decision-makers to adopt leading‑edge observability and reliability technologies.
  • We contribute to our culture of diversity, opportunity, inclusion, and respect.
Technologies
  • AWS
  • OpenSearch
  • Cloud
  • Datadog
  • Docker
  • Dynatrace
  • ElasticSearch
  • GitLab
  • Grafana
  • Support
  • Java
  • Jenkins
  • Kubernetes
  • Linux
  • Load Balancing
  • OpenTelemetry
  • Prometheus
  • Python
  • Security
  • Splunk
  • TCP/IP
  • Terraform
  • DevOps
More

We are J.P. Morgan Chase, a global leader in financial services and a leader across banking, markets, securities services, and payments through our Commercial & Investment Bank. We provide strategic advice and products to major corporations, governments, wealthy individuals, and institutional investors in more than 100 countries. Our first-class business in a first-class way approach drives everything we do, and we build trusted, long‑term partnerships to help clients achieve their business objectives. We value the diverse talents of our people, are committed to equal opportunity, and foster a culture of diversity, inclusion, and respect. This is a full‑time role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

慨正橡扯 • Glasgow

On-site
GBP 70,000 - 90,000
Senior Lead Site Reliability / DevOps Engineer
Senior Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Lead Site Reliability / DevOps Engineer
Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 130,000
Senior Lead SRE: Reliability, Observability & Resiliency
Senior Lead SRE: Reliability, Observability & Resiliency

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Software Engineer - Java
Lead Software Engineer - Java

JPMorganChase • Glasgow

On-site
GBP 90,000 - 150,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000