Site Reliability Engineer

iXceed Solutions

Basildon

On-site

GBP 55,000 - 75,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A leading technology company in the United Kingdom is seeking a highly skilled Site Reliability Engineer to design, deploy, and support cloud-native applications. Ideal candidates will have expertise in Kubernetes and a strong background in cloud technologies like AWS, Azure, or GCP. Responsibilities include ensuring system reliability, automating processes, and improving performance while collaborating closely with development and operations teams. A Bachelor’s degree in a relevant field is required.

Qualifications

  • 3 years experience as an SRE, DevOps Engineer, or similar role supporting large-scale systems.
  • Hands-on experience with major cloud providers (AWS, Azure, or GCP).
  • Strong knowledge of Linux systems administration and networking concepts.

Responsibilities

  • Design, deploy, and manage Kubernetes clusters.
  • Automate infrastructure provisioning and application deployment.
  • Monitor, troubleshoot, and optimize system performance.

Skills

Kubernetes expertise
Cloud technologies (AWS, Azure, GCP)
Scripting/Programming (Python, Bash, Go)
IaC tools (Terraform, Helm, CloudFormation)
Linux systems administration
Monitoring tools (Prometheus, Grafana)
CI/CD tools (Jenkins, GitLab CI)
Excellent troubleshooting skills

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

Terraform
Prometheus
GitLab CI

Job description

We are seeking a highly skilled Site Reliability Engineer (SRE) with deep expertise in Kubernetes and cloud technologies AWS, Azure, or GCP.

The SRE will be responsible for designing, deploying, automating, and supporting highly available, scalable, and secure containerized applications in cloud-native environments. You will work closely with development, operations, and security teams to ensure the reliability, performance, and efficiency of our production systems.

Key Responsibilities
  • Design, deploy, and manage Kubernetes clusters on‑premises and/or cloud‑managed such as EKS, AKS, GKE to support scalable microservices architectures.
  • Automate infrastructure provisioning and application deployment using Infrastructure as Code (IaC) tools such as Terraform, Helm, or CloudFormation.
  • Monitor, troubleshoot, and optimize system performance using observability tools.
  • Implement and manage CI/CD pipelines to ensure rapid, repeatable, and reliable software delivery.
  • Ensure system reliability, availability, and security through proactive monitoring, incident response, and root cause analysis.
  • Develop and maintain runbooks, dashboards, and documentation for operational procedures and system architectures.
  • Participate in on‑call rotations and respond to production incidents, ensuring minimal downtime and fast recovery.
  • Collaborate with development and operations teams to drive DevOps and SRE best practices including capacity planning, scaling, and cost optimization.
  • Continuously improve automation tooling and processes to reduce manual work and increase system reliability.
Required Skills & Experience
  • 3 years experience as an SRE, DevOps Engineer, or similar role supporting large‑scale systems.
  • Expertise in Kubernetes deployment, scaling, upgrades, troubleshooting, and networking.
  • Hands‑on experience with at least one major cloud provider (AWS, Azure, or GCP).
  • Proficiency in scripting/programming (Python, Bash, Go, etc.).
  • Experience with IaC tools (Terraform, Helm, CloudFormation, ARM, etc.).
  • Strong knowledge of Linux systems administration and networking concepts.
  • Familiarity with monitoring, logging, and ingestion tools (Prometheus, Grafana, ELK, EFK).
  • Experience with CI/CD tools (Jenkins, GitLab CI, ArgoCD, etc.).
  • Understanding of security best practices in cloud and containerized environments.
  • Excellent troubleshooting and problem‑solving skills.
  • Strong communication and collaboration skills.
Preferred Qualifications
  • Certified Kubernetes Administrator (CKA) or similar certification.
  • Experience with service mesh (Istio, Linkerd), ingress controllers, and API gateways.
  • Experience in a multicloud or hybrid cloud environment.
  • Familiarity with GitOps practices and tools (ArgoCD, Flux).
  • Experience with disaster recovery, backup, and business continuity planning.
Education

Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

This role is ideal for engineers who are passionate about automation, reliability, and modern cloud‑native architectures and who thrive in fast‑paced, collaborative environments.

We are an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, protected veteran status, or disability status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) - Cloud Kubernetes Platform
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform

Intuition IT Solutions Ltd • Glasgow

Hybrid
GBP 60,000 - 75,000
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Kubernetes SRE: Cloud-Native Reliability & Automation
Kubernetes SRE: Cloud-Native Reliability & Automation

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Factset • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage
Free lunch in the office (Mon–Fri)
Employee social events and sports
SRE
SRE

Technopride Ltd • Hove

On-site
GBP 60,000 - 80,000
Site Reliability Engineer (SRE) – Cloud Platforms
Site Reliability Engineer (SRE) – Cloud Platforms

Talenzon group • Greater London

On-site
GBP 70,000 - 110,000
Senior Site Reliability Engineer (Platform Reliability, Resilience)
Senior Site Reliability Engineer (Platform Reliability, Resilience)

Elastic • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage for you and family
Flexible location & schedule
Generous vacation days
+3
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Site Reliability Engineer (DV Security Clearance)
Site Reliability Engineer (DV Security Clearance)

Onyx-Conseil • Manchester

On-site
GBP 90,000 - 120,000