Senior Site Reliability Engineer – Cloud, Security & Automation

Hackajob Ltd

Leeds

On-site

GBP 61,000 - 101,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Hackajob Ltd. is seeking a Site Reliability Engineer to enhance the reliability, observability, and operability of customer‑facing digital banking services. You will join a distributed, cross‑functional team to reduce toil and improve resilience across microservices.

The role requires hands-on experience with Kubernetes, AWS, and monitoring stacks (Grafana/Prometheus/Elasticsearch/Kibana). Strong scripting and AI-assisted tooling experience are also valued.

Qualifications

  • Formal training or certification in software engineering concepts.
  • Proven software engineering experience in at least one language (Python/Go/Java).
  • Experience designing, coding, testing and delivering software in at least one tech stack.
  • Strong debugging and troubleshooting across distributed systems.
  • Experience as an SRE or supporting production services in an SRE capacity.
  • Working knowledge of microservice infrastructure: service discovery, ingress, networking, load balancing.
  • Experience with Kubernetes and cloud computing services.
  • Familiarity with observability tools (Grafana, Prometheus, Elasticsearch, Kibana, Jaeger).
  • Ability to use AI-assisted engineering tools responsibly, validate outputs, and handle sensitive information securely.
  • Preferred: AWS experience.
  • Preferred: building internal reliability tooling (CLI tools, automation pipelines, operators).
  • Preferred: improving developer experience via golden paths, templates, or reusable patterns.
  • Preferred: applying AI to operational workflows with approved tools and patterns.

Responsibilities

  • Improve reliability, monitoring, and alerting of mission-critical microservices.
  • Reduce operational toil by automating processes and building reliable infra and tooling.
  • Develop service metrics, SLIs/SLOs, error budgets, dashboards, and alerts.
  • Work with development teams to design services for reliability and scale.
  • Design self-healing and resiliency patterns (graceful degradation, rate limiting, circuit breakers, failover).
  • Partner with engineering, product, and platform teams to promote reliability standards.
  • Conduct performance testing and capacity planning to address bottlenecks.
  • Incorporate metrics, alerting, logging, automation, resiliency, capacity in feature planning.
  • Use approved AI tools for root-cause analysis, log/traces, runbooks, post-incident reviews, test scaffolding, docs.
  • Develop AI skills: prompting, output validation, automation workflows, safe usage practices.

Skills

Python
Java
Go
Kubernetes
AWS
Prometheus
Grafana
Elasticsearch
Kibana
Service discovery
Load balancing
Cloud
Observability
AI tooling
Distributed systems
Microservices
Networking

Education

Software engineering certification

Tools

Kubernetes
Prometheus
Grafana
Elasticsearch
Kibana
AWS

Job description

Hackajob Ltd. is seeking a Site Reliability Engineer to enhance the reliability, observability, and operability of customer‑facing digital banking services. You will join a distributed, cross‑functional team to reduce toil and improve resilience across microservices.

The role requires hands-on experience with Kubernetes, AWS, and monitoring stacks (Grafana/Prometheus/Elasticsearch/Kibana). Strong scripting and AI-assisted tooling experience are also valued.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems & SRE Manager - Uptime & Automation
Senior Systems & SRE Manager - Uptime & Automation

Hackajob Ltd • Leeds

On-site
GBP 61,000 - 101,000
Lead Site Reliability Engineer: Architect & Incident Leader
Lead Site Reliability Engineer: Architect & Incident Leader

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Cloud Security Platform Engineer: Kubernetes & Cloud-Native
Cloud Security Platform Engineer: Kubernetes & Cloud-Native

Hackajob Ltd • Leeds

On-site
GBP 61,000 - 101,000
On-site work
Office in Wiltshire or London
On-call rotation
Senior Site Reliability Engineer - Kubernetes & Automation
Senior Site Reliability Engineer - Kubernetes & Automation

20035 FactSet Europe Limited • Greater London

Hybrid
GBP 90,000 - 150,000
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 65,000 - 120,000
Life Insurance 4x Annual Salary
Private Medical Insurance
Bonus Scheme
+3
Senior SRE - AWS Platform & Reliability Lead
Senior SRE - AWS Platform & Reliability Lead

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
SRE Engineer – FinTech Reliability, Observability & Cloud
SRE Engineer – FinTech Reliability, Observability & Cloud

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 90,000 - 130,000
Global engineering organisation
Engineering-led culture
Technically challenging problems
+1
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Senior Platform Engineer: AI-Driven Cloud & Kubernetes
Senior Platform Engineer: AI-Driven Cloud & Kubernetes

Hackajob Ltd • Leeds

On-site
GBP 90,000 - 110,000
Senior Site Reliability Engineer — Cloud & Observability
Senior Site Reliability Engineer — Cloud & Observability

GCA Altium • Cambridge

On-site
GBP 90,000 - 140,000