Senior Site Reliability Engineer (Application / API Focused)

Xpertise Recruitment

Greater London

Hybrid

GBP 90,000 - 110,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

25% Bonus
Excellent Benefits

Job summary

A leading recruitment firm is looking for a Senior Site Reliability Engineer specializing in application and API reliability. This hybrid role involves working closely with product and engineering teams to enhance the reliability of customer-facing platforms. Candidates should have a strong background in digital product environments and expertise in API architectures, microservices, and observability tools. Excellent salary and benefits package offered.

Qualifications

  • Proven experience operating as an SRE within digital product environments.
  • Strong understanding of API architectures and microservices.
  • Hands-on experience defining SLIs, SLOs, and error budgets.

Responsibilities

  • Embed SRE practices across API and microservices-based architectures.
  • Define and own meaningful SLIs/SLOs aligned to customer journeys.
  • Improve service reliability through proactive observability and alert tuning.

Skills

API architectures
Microservices
Distributed systems behaviour
Observability tools (Datadog, Splunk, Prometheus)
Incident response
Kubernetes

Tools

Datadog
Splunk
Prometheus

Job description

Senior Site Reliability Engineer (Application / API Focused)

Location: London (Hybrid)

Salary: £100,000 per annum + 25% Bonus + Excellent Benefits

We are hiring a Senior SRE to support a large-scale digital organisation undergoing a major commercial re-platforming across web and mobile channels.

This role sits much closer to the application layer than traditional infrastructure SRE positions. You will work directly with product and engineering teams across customer-facing platforms (web, mobile, payment journeys, APIs) to improve reliability, resilience, and service behaviour in production.

This is not a ticket-driven operational role and not a pure platform engineering post. It is about embedding measurable reliability into distributed systems at service level.

What You’ll Be Doing
  • Embed SRE practices across API and microservices-based architectures
  • Define and own meaningful SLIs/SLOs aligned to customer journeys and business-critical flows
  • Improve service reliability through proactive observability, tracing, telemetry and alert tuning
  • Partner closely with backend and platform engineers to reduce systemic failure modes
  • Lead and contribute to incident response, post-incident reviews and resilience improvements
  • Move the organisation from symptom-based alerting to customer-impact driven diagnostics
  • Contribute to release safety, progressive deployments and production guardrails
What We’re Looking For
  • Proven experience operating as an SRE within digital product environments
  • Strong understanding of API architectures, microservices and distributed systems behaviour
  • Hands-on experience defining and implementing SLIs, SLOs and error budgets
  • Deep observability exposure (e.g. Datadog, Splunk, Prometheus, tracing/APM platforms)
  • Experience working closely with application engineering teams, not just infrastructure teams
  • Background in high-availability, customer-facing systems where outages have commercial impact
  • Cloud-native exposure (AWS preferred) with practical understanding of Kubernetes environments
Important

This role is best suited to engineers who care deeply about production behaviour, customer experience in failure scenarios, and reliability as a first-class product feature, rather than engineers focused purely on infrastructure provisioning or CI/CD enablement.

Please get in touch with Benjamin Applewhaite to discuss the role in confidence.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Senior Site Reliability Engineer (LON)
Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd • Greater London

Hybrid
GBP 90,000 - 150,000
Benefits as Cash
Hybrid work model
SRE Technical Lead
SRE Technical Lead

83zero Ltd • United Kingdom

Hybrid
GBP 90,000 - 110,000
Salary up to 100,000
5% annual bonus
Hybrid working model
+1
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Gravitas Recruitment Group (Global) Ltd • Greater London

On-site
GBP 75,000 - 100,000
Site Reliability Engineer (remote working)
Site Reliability Engineer (remote working)

Vertus Partners • Greater London

Hybrid
GBP 77,000 - 104,000
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

ScaleneWorks People Solutions LLP • Bournemouth

On-site
GBP 60,000 - 80,000
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

On-site
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14