Site Reliability Engineer

Queen Square Recruitment Ltd

Greater London

On-site

GBP 90,000 - 120,000

Full time

7 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Queen Square Recruitment Ltd is seeking an experienced Site Reliability Engineer (SRE) for a 1 year contract. The role focuses on reliability, availability, scalability and performance of production systems on Google Cloud Platform (GCP).

You will work across cloud infrastructure, containers, automation, observability and incident management to help engineering teams build secure, resilient services. The ideal candidate will have strong hands-on SRE experience with GCP, Kubernetes, Docker,

Qualifications

  • Experience in designing, building and operating scalable, reliable systems in production.
  • Strong hands-on SRE experience with Google Cloud Platform (GCP) and containerised apps.
  • Proficiency in scripting languages such as Python, Java or Bash.

Responsibilities

  • Define, implement and manage SLIs and SLOs.
  • Develop monitoring dashboards, alerting and observability capabilities.
  • Participate in incident response, RCA and permanent corrective actions.
  • Develop automation to reduce manual operational effort.
  • Automate runbooks, recovery procedures and routine support activities.
  • Implement monitoring, logging, metrics and telemetry for production visibility.
  • Improve availability, scalability, resilience, security and performance of systems.
  • Support and troubleshoot GCP cloud environments.
  • Work with Docker and Kubernetes to manage containerised apps.
  • Share SRE best practices across teams.

Skills

GCP
Kubernetes
Docker
Linux/Unix
Python
Java
Bash
SRE
observability
incident management
automation
SLIs/SLOs
monitoring
logging
Terraform
GKE
Prometheus
Grafana

Tools

Docker
Kubernetes
Terraform
Prometheus
Grafana
Datadog
Splunk
Cloud Monitoring

Job description

1 year Contract

Our client is looking for an experienced Site Reliability Engineer (SRE) to join an engineering team within a major financial services environment. This is a hands-on engineering role focused on ensuring the reliability, availability, scalability and performance of production systems running on Google Cloud Platform (GCP).

You will work across cloud infrastructure, containerised applications, automation, observability and incident management, helping engineering teams build and operate secure, resilient and highly available services. The ideal candidate will have strong experience with GCP, Kubernetes, Docker, Linux/Unix and scripting using Python, Java or Bash, alongside practical experience improving production reliability through automation and monitoring.

Key Responsibilities
  • Define, implement and manage Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Develop meaningful monitoring dashboards, alerting and observability capabilities.
  • Participate in incident response, root-cause analysis (RCA) and permanent corrective actions.
  • Develop automation using Python, Java and/or Bash to reduce manual operational effort.
  • Automate operational runbooks, recovery procedures and routine support activities.
  • Implement monitoring, logging, metrics and telemetry to improve production visibility.
  • Improve application and infrastructure availability, scalability, resilience, security and performance.
  • Support and troubleshoot GCP cloud environments.
  • Work with Docker and Kubernetes to manage containerised applications.
  • Collaborate with application, platform, cloud, security and infrastructure engineering teams.
  • Improve production readiness, deployment reliability, rollback and release processes.
  • Share technical knowledge and promote SRE best practices across engineering teams.
  • Strong hands-on Site Reliability Engineering / Production Engineering experience.
  • Experience working with Google Cloud Platform (GCP).
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with Docker and Kubernetes.
  • Strong scripting/coding experience using Python, Java and/or Bash.
  • Good understanding of networking protocols and troubleshooting.
  • Experience with monitoring, logging, metrics and observability.
  • Knowledge of SLIs, SLOs, alerting and service reliability principles.
  • Experience with incident management, root-cause analysis and production troubleshooting.
  • Experience building automation to improve operational efficiency and system reliability.
  • Understanding of scalable, highly available and resilient system design.
  • Experience with Google Kubernetes Engine (GKE).
  • Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk or Cloud Monitoring.
  • Infrastructure as Code using Terraform.
  • CI/CD and automated deployment pipelines.
  • Experience with incident response, post-incident reviews and reliability improvements.
  • Experience working in Banking / Financial Services.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GCP SRE Engineer: Reliability & Automation (Contract)
GCP SRE Engineer: Reliability & Automation (Contract)

Queen Square Recruitment Ltd • Greater London

On-site
GBP 90,000 - 120,000
SRE Engineer
SRE Engineer

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Mphasis • Greater London

On-site
GBP 90,000 - 120,000
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform

Intuition IT Solutions Ltd • Glasgow

Hybrid
GBP 60,000 - 75,000
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Devops SRE
Devops SRE

Test Triangle • Greater London

On-site
GBP 70,000 - 90,000
Site Reliability Engineer (SRE) – Cloud Platforms
Site Reliability Engineer (SRE) – Cloud Platforms

Talenzon group • Greater London

On-site
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Site Reliability Engineer (SRE) – Cloud Engineer
Site Reliability Engineer (SRE) – Cloud Engineer

Dianaduggan • Glasgow

On-site
GBP 74,000 - 83,000