Senior Site Reliability Engineer

P2P

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

P2P is seeking a Senior Site Reliability Engineer to join our agile engineering team in London. You will drive reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership.

You will design, implement, and operate cloud-native infrastructure on GCP using Kubernetes, Terraform, and Helm, while championing automation, observability, and best practices across the SDLC.

Qualifications

  • 4+ years of hands-on experience with Google Cloud Platform (GCP) or similar cloud infrastructure.
  • Expert-level proficiency in managing, scaling, and troubleshooting production Kubernetes environments.
  • Deep expertise in Terraform for managing cloud and Kubernetes resources.
  • Strong experience with Helm for packaging and deploying applications on Kubernetes.
  • Proficient in Python for automation and tool development.
  • Strong command-line skills with Linux systems.

Responsibilities

  • Architect, deploy, and maintain scalable infrastructure on GCP using Kubernetes and Infrastructure-as-Code tools.
  • Champion automation across the entire SDLC, utilizing IaC, Python and Bash to reduce toil and improve efficiency.
  • Own and evolve declarative infrastructure using Terraform and Helm for cloud resources and Kubernetes deployments.
  • Implement and manage robust monitoring, alerting, and logging solutions for clear system visibility.
  • Define, measure, and enforce SLOs/SLIs; participate in on-call rotation and post-incident reviews.
  • Collaborate with software development teams to advise on deployment strategies and cloud-native practices.
  • Take full ownership of projects from inception through production operation, including documentation transfer.

Skills

High Ownership
Small Team Mentality
Adaptability
Communication

Tools

GCP
Kubernetes
Terraform
Helm
Python
Bash
CI/CD pipelines
Linux

Job description

Senior Site Reliability Engineer (SRE) - GCP/Kubernetes

We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership.

The successful candidate will split their time between providing expert operational support for our critical systems and leading exciting new infrastructure projects. Our mindset is to get the right person not the person with the skills that match our stack. It is important to be able to foresee problems before they show up and create solutions that mitigate them. If you enjoy a challenging environment, implementing “infrastructure as code” principles, and directly seeing the impact of your work, this is the place for you.

What You Will Do
  • Design & Build: Architect, deploy, and maintain highly scalable and reliable infrastructure on Google Cloud Platform (GCP) using Kubernetes and Infrastructure-as-Code tools.
  • Automation: Champion automation across the entire software development lifecycle (SDLC), utilizing IaC, Python and Bash to reduce toil and improve operational efficiency.
  • Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment.
  • Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification.
  • Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post-incident reviews to drive continuous improvement.
  • Collaboration: Work closely with software development teams to provide expert guidance on deployment strategies, scalability concerns, and cloud-native best practices.
  • Ownership: Take full ownership of projects from inception through to production operation, including documentation and knowledge transfer.
Required Experience & Skills
Core Technical Stack
  • Cloud Platform: 4+ years of hands‑on experience with Google Cloud Platform (GCP) (or similar cloud infrastructure).
  • Container Orchestration: Expert‑level proficiency in managing, scaling, and troubleshooting production Kubernetes environments.
  • Infrastructure-as-Code: Deep expertise in Terraform for managing cloud and Kubernetes resources.
  • Deployment: Strong experience with Helm for packaging and deploying applications on Kubernetes.
  • Scripting/Programming: Proficient in at least one major programming language, preferably Python, for automation and tool development.
Tooling & Concepts
  • CI/CD: Experience setting up and maintaining modern CI/CD pipelines.
  • Observability: Practical experience implementing and managing monitoring and logging tools.
  • Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud‑native networking within Kubernetes.
  • Operating Systems: Strong command‑line skills and experience with Linux systems.
Soft Skills & Team Fit
  • High Ownership: Demonstrated ability to own a problem end‑to‑end, from investigation to resolution and preventative measures.
  • Small Team Mentality: Happy to be a generalist and switch context quickly between support tickets, operational toil reduction, and long‑term project work.
  • Adaptability: A proven track record of rapidly learning and applying new technologies and tools. Equivalent experience with other clouds (AWS/Azure) or similar tools is highly valued.
  • Communication: Excellent verbal and written communication skills for documentation and interacting with non‑technical stakeholders.
Bonus Points For
  • Familiarity with Service Mesh technologies (e.g., Istio).
  • Experience in security best practices within cloud and container environments (e.g., hardening, secrets management).
  • Certifications in GCP or Kubernetes (e.g., CKAD, CKA, Professional Cloud DevOps Engineer).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
Site Reliability Engineer-- API Integration
Site Reliability Engineer-- API Integration

DELTACLASS TECHNOLOGY SOLUTIONS LIMITED • Leeds

On-site
GBP 60,000 - 90,000
Devops SRE
Devops SRE

Test Triangle • Greater London

On-site
GBP 70,000 - 90,000
Site Reliability Engineer (SRE) / Platform Engineer
Site Reliability Engineer (SRE) / Platform Engineer

Adecco • City Of London

Hybrid
Hybrid work arrangement
Competitive day rate
London-based contract
Senior SRE: GCP, Kubernetes & Automation Leader
Senior SRE: GCP, Kubernetes & Automation Leader

Brevan Howard CFD LTD • Greater London

On-site
GBP 70,000 - 95,000
GCP Cloud Engineer
GCP Cloud Engineer

DELTACLASS TECHNOLOGY SOLUTIONS LIMITED • Leeds

On-site
GBP 70,000 - 90,000
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Senior Site Reliability Engineer - Selby Jennings
Senior Site Reliability Engineer - Selby Jennings

eFinancialCareers • Greater London

On-site
GBP 90,000 - 130,000
Senior SRE: GCP & Kubernetes, Automation Lead
Senior SRE: GCP & Kubernetes, Automation Lead

P2P • Greater London

On-site
GBP 90,000 - 130,000
Senior Cloud SRE - Kubernetes, GCP & CI/CD
Senior Cloud SRE - Kubernetes, GCP & CI/CD

Test Triangle • Greater London

On-site
GBP 70,000 - 90,000