Senior Site Reliability Lead: Resiliency & Telemetry Architect

Next Frontier Capital

Glasgow

On-site

GBP 90,000 - 140,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorgan Chase is seeking a Software Engineer in a critical site reliability leadership role within the Commercial & Investment Bank. You will guide resiliency design reviews, mentor engineers, and lead initiatives to improve reliability for large-scale systems.

You will champion observability across OpenTelemetry pipelines, manage on-prem/cloud deployments, and collaborate with stakeholders to define SLOs, error budgets, and incident response playbooks.

Qualifications

  • Formal training or certification on software engineering concepts and advanced applied experience.
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform.
  • Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.)
  • Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc.
  • Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.)
  • Experience with container and container orchestration (e.g., ECS, Kubernetes, Docker, etc.)
  • Hands-on experience with the design, deployment, and operation of OpenTelemetry collectors in production environments, focusing on technical aspects such as configuring, optimizing, and troubleshooting OTLP endpoints and receivers.
  • Ability to expand and collaborate across different levels and stakeholder groups

Responsibilities

  • Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team
  • Leads initiatives to improve the reliability and stability of your team’s applications and platforms using data-driven analytics to improve service levels
  • Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers
  • Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology-related bottlenecks in your areas of expertise
  • Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses
  • Documents and shares knowledge within your organization via internal forums and communities of practice
  • Design, implement, and maintain operational reliability for large-scale Opentelemetry pipelines on hybrid on-prem/cloud environments.
  • Support telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheuse, Elasticsearch, OpenSearch optimizing for performance, real-time monitoring, logging, and alerting.
  • Assess, refactor, and incrementally migrate custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.

Skills

Reliability engineering
Observability
Kubernetes
OpenTelemetry
Containerization
SRE principles
Java
Python
Go

Tools

Grafana
Prometheus
Datadog
Elasticsearch
OpenTelemetry
Terraform
Docker
Kubernetes

Job description

JPMorgan Chase is seeking a Software Engineer in a critical site reliability leadership role within the Commercial & Investment Bank. You will guide resiliency design reviews, mentor engineers, and lead initiatives to improve reliability for large-scale systems.

You will champion observability across OpenTelemetry pipelines, manage on-prem/cloud deployments, and collaborate with stakeholders to define SLOs, error budgets, and incident response playbooks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer - Observability & Resilience
Lead Site Reliability Engineer - Observability & Resilience

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 120,000
Lead Site Reliability Engineer: Architect Resilient Systems
Lead Site Reliability Engineer: Architect Resilient Systems

Hackajob Ltd • Milton

On-site
GBP 110,000 - 140,000
Senior Site Reliability & DevOps Leader
Senior Site Reliability & DevOps Leader

JPMorgan Chase & Co. • Glasgow

On-site
GBP 85,000 - 110,000
Senior Site Reliability & DevOps Lead
Senior Site Reliability & DevOps Lead

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 130,000
Lead Site Reliability Engineer — SRE Leadership
Lead Site Reliability Engineer — SRE Leadership

Hackajob Ltd • Cumbernauld

Hybrid
GBP 85,000 - 130,000
Lead Site Reliability Engineer - Resilience & Observability
Lead Site Reliability Engineer - Resilience & Observability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 150,000
Senior Lead Site Reliability / DevOps Engineer
Senior Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. MORGAN • Glasgow

On-site
GBP 90,000 - 120,000
Senior Lead SRE: Reliability, Observability & Resiliency
Senior Lead SRE: Reliability, Observability & Resiliency

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 75,000 - 100,000
Lead Site Reliability / DevOps Engineer
Lead Site Reliability / DevOps Engineer

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 130,000