Lead SRE - Charing Cross

Hackajob Ltd

Leeds

On-site

GBP 61,000 - 101,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Hackajob Ltd. is seeking a Site Reliability Engineer to enhance the reliability, observability, and operability of customer‑facing digital banking services. You will join a distributed, cross‑functional team to reduce toil and improve resilience across microservices.

The role requires hands-on experience with Kubernetes, AWS, and monitoring stacks (Grafana/Prometheus/Elasticsearch/Kibana). Strong scripting and AI-assisted tooling experience are also valued.

Qualifications

  • Formal training or certification in software engineering concepts.
  • Proven software engineering experience in at least one language (Python/Go/Java).
  • Experience designing, coding, testing and delivering software in at least one tech stack.
  • Strong debugging and troubleshooting across distributed systems.
  • Experience as an SRE or supporting production services in an SRE capacity.
  • Working knowledge of microservice infrastructure: service discovery, ingress, networking, load balancing.
  • Experience with Kubernetes and cloud computing services.
  • Familiarity with observability tools (Grafana, Prometheus, Elasticsearch, Kibana, Jaeger).
  • Ability to use AI-assisted engineering tools responsibly, validate outputs, and handle sensitive information securely.
  • Preferred: AWS experience.
  • Preferred: building internal reliability tooling (CLI tools, automation pipelines, operators).
  • Preferred: improving developer experience via golden paths, templates, or reusable patterns.
  • Preferred: applying AI to operational workflows with approved tools and patterns.

Responsibilities

  • Improve reliability, monitoring, and alerting of mission-critical microservices.
  • Reduce operational toil by automating processes and building reliable infra and tooling.
  • Develop service metrics, SLIs/SLOs, error budgets, dashboards, and alerts.
  • Work with development teams to design services for reliability and scale.
  • Design self-healing and resiliency patterns (graceful degradation, rate limiting, circuit breakers, failover).
  • Partner with engineering, product, and platform teams to promote reliability standards.
  • Conduct performance testing and capacity planning to address bottlenecks.
  • Incorporate metrics, alerting, logging, automation, resiliency, capacity in feature planning.
  • Use approved AI tools for root-cause analysis, log/traces, runbooks, post-incident reviews, test scaffolding, docs.
  • Develop AI skills: prompting, output validation, automation workflows, safe usage practices.

Skills

Python
Java
Go
Kubernetes
AWS
Prometheus
Grafana
Elasticsearch
Kibana
Service discovery
Load balancing
Cloud
Observability
AI tooling
Distributed systems
Microservices
Networking

Education

Software engineering certification

Tools

Kubernetes
Prometheus
Grafana
Elasticsearch
Kibana
AWS

Job description

Salary: £61,000 - 101,000 per year

Requirements:
  • Formal training or certification in software engineering concepts, with advanced applied experience.
  • Proven software engineering experience and proficiency in at least one programming language, such as Python, Go, or Java.
  • Experience designing, coding, testing, and delivering software in at least one technology stack.
  • Strong debugging and troubleshooting skills across distributed systems.
  • Experience as a Site Reliability Engineer or supporting production services in an SRE capacity.
  • Working knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing.
  • Experience with Kubernetes and cloud computing services.
  • Familiarity with observability and reliability tools such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger.
  • Ability to use AI-assisted engineering tools responsibly, validate outputs, understand failure modes, and handle sensitive information securely.
  • Preferred: Experience with AWS.
  • Preferred: Experience building internal reliability tooling, such as command-line tools, automation pipelines, operators, or controllers.
  • Preferred: Experience improving developer experience through golden paths, paved roads, templates, or reusable engineering patterns.
  • Preferred: Experience applying AI to operational workflows, such as alert enrichment, summarisation, runbook generation, or anomaly triage, using approved tools and patterns.
Responsibilities:
  • Improve the reliability, monitoring, and alerting of mission-critical microservices.
  • Reduce operational toil by automating processes and building reliable infrastructure and tooling that accelerates feature development.
  • Develop service metrics, user journeys, service-level indicators and objectives, error budgets, dashboards, and actionable alerts.
  • Work with development teams throughout the software lifecycle to design services for reliability and scale.
  • Design and implement self-healing and resiliency patterns, including graceful degradation, rate limiting, circuit breakers, and failover strategies.
  • Partner with engineering, product, and platform teams to promote reliability standards and adoption.
  • Conduct performance testing and capacity planning to identify and address bottlenecks proactively.
  • Participate in feature planning to incorporate metrics, alerting, logging, automation, resiliency, capacity, and performance needs from the outset.
  • Use approved AI tools to support root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation.
  • Develop role-relevant AI skills, including effective prompting, output validation, automation workflows, and safe usage practices.
Technologies:
  • AI
  • AWS
  • Cloud
  • ElasticSearch
  • Grafana
  • Support
  • Java
  • Kibana
  • Kubernetes
  • Load Balancing
  • Prometheus
  • Python
  • microservices
  • Network

We are a global financial services leader serving prominent corporations, governments, wealthy individuals, and institutional investors, with a focus on trusted, long-term client partnerships. Our International Consumer Bank is expanding from the US into the UK and Europe, transforming digital banking through intuitive customer experiences. As a Site Reliability Engineer, you will join a diverse, inclusive, geographically distributed team working to improve the reliability, resilience, observability, and operability of customer-facing digital banking services. Our Corporate Technology team develops applications and provides technology support across functions including Finance, Treasury, Risk Management, Human Resources, Compliance, and Legal, while supporting evolving technology needs and controls.

last updated 40 week of 2026

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 80,000 - 100,000
Healthcare
Retirement Planning
Paid Volunteering Days
+1
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

On-site
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

National Health Service • Greater London

Hybrid
GBP 42,000 - 52,000
Hybrid working
Flexible working arrangements
On-site core HQs
Senior Site Reliability Engineer (Application / API Focused)
Senior Site Reliability Engineer (Application / API Focused)

Xpertise Recruitment • Greater London

On-site
GBP 90,000 - 110,000
25% Bonus
Excellent Benefits
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
Systems Engineering Manager - Charing Cross
Systems Engineering Manager - Charing Cross

Hackajob Ltd • Leeds

On-site
GBP 61,000 - 101,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

London Stock Exchange Group • Nottingham

On-site
GBP 90,000 - 120,000
Healthcare
Retirement planning
Paid volunteering days
+1
Senior Site Reliability Engineer (LON)
Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd • Greater London

Hybrid
GBP 90,000 - 150,000
Benefits as Cash
Hybrid work model
Lead Site Reliability Engineer - Glasgow
Lead Site Reliability Engineer - Glasgow

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000