Site Reliability Engineer

Dabster

Leeds

On-site

GBP 65,000 - 85,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Dabster in Leeds is seeking a Senior Site Reliability Engineer to lead production systems on Kubernetes (GKE) across public cloud environments. You will architect and maintain CI/CD pipelines, IaC with Terraform, and drive reliability through monitoring and incident response.

The role emphasizes on-call rotations, performance tuning, and collaboration with development and operations teams to automate deploys and improve fault tolerance using Prometheus, Grafana, and OpenTelemetry.

Qualifications

  • 4–8 years of Site Reliability Engineering, DevOps, or Cloud Infrastructure experience.
  • Hands-on production experience with Kubernetes clusters on GKE.
  • Experience with Terraform and automation of cloud infrastructure.
  • CI/CD pipeline management using Jenkins.
  • Solid understanding of SRE principles (incident management, blameless postmortems, capacity planning, and error budgets).
  • Proficiency in Python or similar scripting languages.
  • Strong analytical, troubleshooting, and problem-solving abilities.
  • Excellent written and verbal communication skills.
  • Experience with observability tools such as Prometheus and Grafana (OpenTelemetry).
  • Exposure to GitOps tools like Flux.

Responsibilities

  • Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (GKE).
  • Implement and maintain infrastructure as code using Terraform.
  • Collaborate with development and operations teams to improve system reliability and deployment automation.
  • Build and maintain CI/CD pipelines using Jenkins or similar tools.
  • Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.
  • Automate operational tasks using Python or other scripting languages.
  • Contribute to observability and monitoring improvements using modern tools and best practices.
  • Participate in on-call rotations and incident response processes.

Skills

SRE principles
Incident management
Capacity planning
Error budgets
Python scripting
Troubleshooting
Communication
Blameless postmortems

Tools

Terraform
Kubernetes (GKE)
Jenkins
Prometheus
Grafana
OpenTelemetry
Flux (GitOps)

Job description


  • Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Google GKE).

  • Implement and maintain infrastructure as code using tools such as Terraform.

  • Collaborate with development and operations teams to improve system reliability and deployment automation.

  • Build and maintain CI/CD pipelines using Jenkins or similar tools.

  • Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.

  • Automate operational tasks using Python or other scripting languages.

  • Contribute to observability and monitoring improvements using modern tools and best practices.

  • Participate in on-call rotations and incident response processes.


Required Skills and Experience:


  • 4-8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.

  • Strong hands-on experience managing Kubernetes clusters in production (GKE).

  • Proficiency with Terraform and cloud infrastructure automation.

  • Practical experience with Jenkins and CI/CD pipeline management.

  • Sound understanding of SRE principles (incident management, blameless postmortems, capacity planning, error budgets, etc.).

  • Good programming or scripting skills in Python (preferred) or similar languages.

  • Strong analytical, troubleshooting, and problem-solving abilities.

  • Excellent written and verbal communication skills.

  • Experience with Prometheus, Grafana, or OpenTelemetry for observability.

  • Exposure to GitOps practices and tools (e.g. Flux).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) - Cloud Kubernetes Platform
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform

Intuition IT Solutions Ltd • Glasgow

Hybrid
GBP 60,000 - 75,000
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brevan Howard • Greater London

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

P2P • Greater London

On-site
GBP 90,000 - 130,000
Senior DevOps & Cloud SRE Lead (AWS / K8s)
Senior DevOps & Cloud SRE Lead (AWS / K8s)

GenixBit Labs Pvt. Ltd. • Greater London

Hybrid
GBP 70,000 - 120,000
Senior DevOps & Cloud SRE Lead (AWS / K8s)
Senior DevOps & Cloud SRE Lead (AWS / K8s)

Genixbit • Greater London

Hybrid
GBP 65,000 - 100,000
Kubernetes SRE & GKE Cloud Automation Engineer
Kubernetes SRE & GKE Cloud Automation Engineer

Dabster • Leeds

On-site
GBP 65,000 - 85,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Flowcode • York and North Yorkshire

On-site
GBP 90,000 - 120,000
Unlimited Vacation
Health benefits
Stock options
+6
Site Reliability Engineer (SRE) – Cloud Platforms
Site Reliability Engineer (SRE) – Cloud Platforms

Talenzon group • Greater London

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

Hybrid
GBP 65,000 - 90,000