Data Platform SRE

Outerlimit

Greater London

Hybrid

GBP 90,000 - 125,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work London
Conference attendance
Training budget
Benefits package

Job summary

Outerlimit is seeking a Senior Site Reliability Engineer to embed within the Data Engineering team. You will own reliability, performance, and operability of the data platform, building observability from the ground up and leading incident response for data outages.

You'll balance feature velocity with system stability, applying SRE practices and capacity planning to help the team ship more reliably and efficiently.

Qualifications

  • Proven experience as SRE or in reliability engineering.
  • Strong cloud and data platform experience, especially with Azure.
  • Hands-on with infrastructure-as-code, observability, and incident response.

Responsibilities

  • Define and maintain SLIs/SLOs for data pipelines and streaming systems.
  • Design and maintain data platform infrastructure using IaC in Azure.
  • Improve observability: metrics, logs, traces, data quality monitoring.
  • Lead incident response and blameless post-mortems for outages.
  • Own on-call rotation and reduce toil and alert noise.
  • Automate deployments, scaling, failover, and recovery procedures.

Skills

SRE
Azure
Python
DevOps
Observability

Tools

Databricks
Kafka
Pyspark
Terraform

Job description

Outerlimit is looking for a Senior Site Reliability Engineer to embed within our Data Engineering team. You'll own the reliability, performance, and operability of the data platform — the pipelines, orchestration, storage, and streaming systems that the wider business depends on for analytics, reporting, and product features. This is a hands‑on engineering role: you'll write infrastructure‑as‑code, build observability into data systems from the ground up, lead incident response for data platform outages, and work directly with data engineers to raise the reliability bar on everything they ship.

You’ll act as the reliability voice inside the Data Engineering team — not a separate, siloed SRE function — balancing feature velocity against system stability, and bringing SRE practice (SLOs, error budgets, blameless post‑mortems, capacity planning) to a team that has historically optimised for delivery speed.

Key Responsibilities
  • Define and maintain SLIs/SLOs for critical data pipelines, warehouses, and streaming systems, and use error budgets to guide the pace of change.
  • Design, build, and maintain the infrastructure that underpins the data platform (compute, storage, orchestration, networking) using infrastructure-as-code primarily focussed in Microsoft Azure.
  • Build and improve observability for data systems — metrics, logging, tracing, and data‑quality/freshness monitoring — so issues are caught before they reach downstream consumers.
  • Lead incident response for data platform issues: triage, coordinate, drive root cause analysis, and run blameless post‑mortems that result in durable fixes.
  • Own on‑call rotation design and participate in on‑call for the data platform, working to reduce toil and alert noise over time.
  • Partner with data engineers on pipeline design reviews, capacity planning, and cost/performance trade-offs, embedding reliability practices into their day‑to‑day workflow rather than gatekeeping after the fact.
  • Automate manual operational work — deployments, scaling, failover, data backfills, recovery procedures — to reduce repetitive load on the team.
  • Drive disaster recovery and business continuity planning for data systems, including backup strategy, failover testing, and documented runbooks.
  • Contribute to the broader platform/infrastructure SRE community at Outerlimit, sharing tooling and practices across teams.
  • Support compliance and audit requirements (e.g. SOC 2, data governance) as they relate to data platform reliability, access control, and change management.
Nice to have
  • Experience with a cloud data warehouse (Databricks, BigQuery, or Redshift) at scale.
  • Experience operating streaming systems (Kafka, Azure Event Hubs, or similar) in production.
  • Good understanding of Python with the ability to read/debug Pyspark jobs and configure tooling.
  • Prior experience formally introducing SRE practices to a team that didn't previously have them.
  • Relevant compliance/security exposure (SOC 2, ISO 27001, or similar frameworks).
What we offer
  • Hybrid working — London office, 2+ days per week, with flexibility around core hours.
  • A genuine seat at the table shaping how the Data Engineering team builds and operates its platform, not a bolt-on ops function.
  • Investment in tooling, training, and conference attendance to keep your SRE practice current.
  • The standard Outerlimit benefits package (details shared during the interview process).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Platform SRE — Hybrid, Obs & Reliability Leader
Senior Data Platform SRE — Hybrid, Obs & Reliability Leader

Outerlimit • Greater London

Hybrid
GBP 90,000 - 125,000
Hybrid work London
Conference attendance
Training budget
+1
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

On-site
GBP 51,000 - 85,000
Bonus
Benefits
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

On-site
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14
Senior Site Reliability Engineer (Platform Reliability, Resilience)
Senior Site Reliability Engineer (Platform Reliability, Resilience)

Elastic • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage for you and family
Flexible location & schedule
Generous vacation days
+3
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
AWS Site Reliability Engineer
AWS Site Reliability Engineer

Marks Sattin • Greater London

On-site
GBP 60,000 - 80,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

London Stock Exchange Group • Nottingham

On-site
GBP 90,000 - 120,000
Healthcare
Retirement planning
Paid volunteering days
+1
Observability SRE
Observability SRE

HCLTech • Greater London

On-site
GBP 70,000 - 95,000
Site Reliability Engineer
Site Reliability Engineer

Biometric Talent Ltd • Manchester

On-site
GBP 40,000 - 65,000
Performance-Based Bonus
Pension Scheme
Hybrid Working
+2