Site Reliability Engineer

The Business Connection Group

Greater London

Remote

GBP 130,000 - 150,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

The Business Connection Group is seeking a Site Reliability Engineer for a remote role in the United Kingdom. You will design the visibility layer that enables reliability for complex distributed systems, including satellite constellations and ground infrastructure.

You will lead observability initiatives, define SLOs/SLIs, and collaborate with cross-functional teams to enforce instrumentation standards while practicing proactive incident response and continuous improvement.

Qualifications

  • 4+ years in SRE/reliability/platform engineering.
  • Hands-on experience with production observability stacks.
  • Experience with GCP and Kubernetes in production.

Responsibilities

  • Design, build and scale a unified observability platform.
  • Define end-to-end SLOs, SLIs and error budgets.
  • Embed instrumentation standards with engineering teams.
  • Automate deployment and lifecycle via Terraform and GitOps.
  • Collaborate with infrastructure for visibility across Kubernetes and multi-cloud.
  • Lead monitoring, incident response and blameless post-mortems.

Skills

Observability for distributed systems
SRE / reliability engineering
Kubernetes
Go
Python
Terraform
GitOps
GCP

Tools

Prometheus
Grafana
Loki
OpenTelemetry
Tempo/Jaeger
ArgoCD

Job description

Site Reliability Engineer | £130,000 - £150,000 | Remote

My client, a pioneering advanced technology organisation — a global leader in software‑defined networking platforms for the aerospace sector.

Their work underpins critical connectivity infrastructure and next‑generation space missions.

This is not a routine “maintain uptime” role — you will design the visibility layer that powers reliability for satellite constellations, ground station networks and deep‑space communications systems.

Key Responsibilities
  • Design, build and scale a unified observability platform (metrics, logging, tracing: Prometheus, Grafana, Loki, OpenTelemetry, Tempo/Jaeger)
  • Define and manage end-to-end SLOs, SLIs and error budgets to ensure production readiness and reliability
  • Partner with engineering to embed standards, set instrumentation best practices, and roll out consistent tooling
  • Automate deployment, scaling and lifecycle management via Infrastructure as Code (Terraform) and GitOps (ArgoCD)
  • Collaborate with infrastructure teams to deliver visibility across Kubernetes and multi-cloud environments
  • Lead monitoring, alerting and incident response; foster proactive reliability, blameless post‑mortems and continuous improvement
What You’ll Bring
  • 4+ years in SRE/reliability/platform engineering — focused on observability for large‑scale distributed systems
  • Hands‑on expertise building, scaling and running production observability stacks; diagnosing complex performance & availability issues
  • Strong production experience with GCP and Kubernetes
  • Practical IaC & GitOps experience for configuration and deployment management
  • Proficient in Go or Python for automation and tooling
  • Proven track record defining, implementing and governing SLO/SLI/error budget frameworks for high‑availability services
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

On-site
GBP 51,000 - 85,000
Bonus
Benefits
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
Senior Site Reliability Engineer (LON)
Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd • Greater London

Hybrid
GBP 90,000 - 150,000
Benefits as Cash
Hybrid work model
Site Reliability Engineer (SRE) – Cloud Platforms
Site Reliability Engineer (SRE) – Cloud Platforms

Talenzon group • Greater London

On-site
GBP 70,000 - 110,000
SRE
SRE

Technopride Ltd • Hove

On-site
GBP 60,000 - 80,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Gravitas Recruitment Group (Global) Ltd • Greater London

On-site
GBP 75,000 - 100,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Understanding Recruitment • United Kingdom

On-site
GBP 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000