Site Reliability Engineer

ReVybe IT Recruitment Limited

Greater London

Hybrid

GBP 51,000 - 85,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus
Benefits

Job summary

ReVybe IT Recruitment Limited is seeking an experienced Site Reliability Engineer for a Central London hybrid role. You will design and operate scalable AWS infrastructure, manage Kubernetes clusters, and drive automation across CI/CD pipelines and observability tools.

The role focuses on reliability, performance, and developer experience, with ownership over platform resilience as the business scales. Hybrid work in Central London is offered with 2–3 days in the office.

Qualifications

  • Commercial experience as an SRE/DevOps or similar role.
  • Strong hands-on AWS expertise.
  • Deep Kubernetes experience.
  • Proficient with Terraform and IaC practices.
  • Experience building CI/CD pipelines (GitHub Actions).
  • Solid observability, metrics, logging, and alerting know-how.
  • Understanding of SLIs/SLOs and incident management.

Responsibilities

  • Design, build, and maintain scalable AWS infrastructure.
  • Manage and optimise Kubernetes environments and workloads.
  • Implement IaC using Terraform and CI/CD pipelines.
  • Develop monitoring, logging, and tracing across the platform.
  • Lead incident response and root cause analysis.
  • Collaborate with software engineers to improve deployment processes.

Skills

AWS
Kubernetes
Terraform
GitHub Actions
Observability
CI/CD
Automation

Tools

GitHub Actions

Job description

Site Reliability Engineer

Central London – Hybrid (2/3 days a week in the office)

Up to £85,000 + Benefits

Build, Scale & Improve the Reliability of a Fast-Growing SaaS Platform

We're partnering with a fast-growing SaaS company that's going through an exciting period of growth and investing heavily in its engineering and platform capabilities.

They're looking for an experienced Site Reliability Engineer (SRE) to join the team and play a key role in building highly reliable, scalable, and observable infrastructure.

This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience.

You'll have genuine ownership and the opportunity to influence how the platform evolves as the business continues to scale.

What You'll Be Doing
  • Design, build, and maintain highly available and scalable AWS infrastructure
  • Manage and optimise Kubernetes environments and containerised workloads
  • Build and maintain infrastructure using Terraform and Infrastructure as Code principles
  • Develop and optimise CI/CD pipelines using GitHub Actions
  • Build and improve comprehensive monitoring and observability across the platform
  • Implement and maintain effective logging, metrics, tracing, alerting, and dashboards
  • Define and improve SLIs, SLOs, and reliability metrics
  • Proactively identify and resolve performance, availability, and reliability issues
  • Lead and contribute to incident response, troubleshooting, and root cause analysis
  • Automate operational processes and eliminate repetitive manual tasks
  • Work closely with software engineers to improve deployment processes, system reliability, and developer experience
  • Help improve platform resilience, scalability, and disaster recovery capabilities
  • Contribute to capacity planning and performance optimisation as the platform scales
  • Establish and champion SRE best practices across the wider engineering function
What We're Looking For
  • Proven commercial experience working as an SRE, DevOps Engineer, Platform Engineer, or similar
  • Strong hands-on experience with AWS
  • Strong experience working with Kubernetes
  • Excellent experience with Terraform and Infrastructure as Code
  • Strong experience building and managing GitHub Actions CI/CD pipelines
  • Solid experience with monitoring and observability tooling
  • Strong understanding of metrics, logging, tracing, alerting, and system health
  • Experience troubleshooting complex production environments
  • Understanding of SLIs, SLOs, SLAs, and error budgets
  • Experience with incident management and root cause analysis
  • Good understanding of cloud networking, security, and infrastructure fundamentals
  • Strong scripting/automation skills
  • A strong understanding of reliability, scalability, performance, and availability
  • Excellent communication skills and the ability to work closely with software engineering teams
  • A proactive mindset and genuine passion for automation and continuous improvement
Don't Tick Every Box?

That's okay.

The company is open to speaking with engineers who may not have experience across every technology listed above.

If you have strong foundations in AWS, Kubernetes, Terraform, and cloud infrastructure, along with a genuine interest in reliability and observability, we'd still love to hear from you.

Why Join?
  • Join a fast-growing SaaS company at an exciting stage of its journey
  • Work with a modern AWS and Kubernetes environment
  • Take ownership of reliability, automation, and platform performance
  • Work with modern observability and monitoring technologies
  • Have genuine influence over engineering and platform decisions
  • Work closely with talented software engineering teams
  • Clear opportunities to progress as the business continues to scale
  • Help shape and mature the company's SRE practices
  • Hybrid working from Central London, 2/3 days per week

If you're an experienced SRE, Platform Engineer or DevOps Engineer who enjoys solving complex reliability challenges and wants to have a real impact within a rapidly growing SaaS business, we'd love to hear from you.

Site Reliability Engineer

Central London | Hybrid – 2/3 days per week

Up to £85,000 + Bonus + Benefits

AWS | Kubernetes | Terraform | GitHub Actions | SRE | Observability | Monitoring | CI/CD | Infrastructure as Code | Reliability | Automation | Cloud

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Technical Lead
SRE Technical Lead

83zero Ltd • Wokingham

Hybrid
GBP 60,000 - 100,000
5% bonus
Hybrid working model
Site Reliability Engineer
Site Reliability Engineer

Wedo Technology Solutions Ltd. • Greater London

Remote
GBP 63,000 - 75,000
Senior Platform Engineer / SRE
Senior Platform Engineer / SRE

Myn • Greater London

Hybrid
GBP 90,000 - 120,000
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Senior AWS Platform Engineer
Senior AWS Platform Engineer

ReVybe IT Recruitment Limited • City Of London

Hybrid
GBP 76,000 - 90,000
Hybrid work in London office (2 days)
Benefits
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000
Site Reliability Engineer SRE Kubernetes
Site Reliability Engineer SRE Kubernetes

Client Server • Cambridge

Hybrid
GBP 59,000 - 81,000
Pension
Private Medical Insurance
Life Assurance
+5
Site Reliability Engineer (SRE) / Platform Engineer
Site Reliability Engineer (SRE) / Platform Engineer

Adecco • City Of London

Hybrid
Hybrid work arrangement
Competitive day rate
London-based contract
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GCS • Glasgow

Hybrid
GBP 75,000 - 110,000