Senior Site Reliability Engineer

Understanding Recruitment

United Kingdom

On-site

GBP 90,000 - 120,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Understanding Recruitment is seeking a Senior Site Reliability Engineer to enhance reliability, observability, and tooling for a latency-sensitive production platform. The role spans production infrastructure, monitoring, incident response, and deployment workflows with a strong Linux and networking focus.

You will work to improve CI/CD, automation, and internal tooling while elevating developer experience from local setups to production.

Qualifications

  • Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering.
  • Experience operating production infrastructure in cloud environments.
  • Strong Linux systems knowledge and understanding of networking fundamentals.
  • Experience with monitoring, observability and alerting.
  • Strong troubleshooting and root cause analysis skills.
  • Experience with CI/CD and infrastructure automation.
  • AWS, Terraform or Ansible experience would be advantageous.
  • Experience with high-performance, high-throughput or latency-sensitive systems would be valuable.

Responsibilities

  • Improve the reliability and operability of production systems.
  • Build and improve monitoring, logging, tracing, dashboards and alerting.
  • Improve incident diagnosis, root cause analysis and operational workflows.
  • Build safer and more repeatable deployment and rollback processes.
  • Automate repetitive operational and infrastructure work.
  • Improve CI/CD pipelines and release processes.
  • Develop internal tooling that helps engineers operate production systems more effectively.
  • Improve the developer experience from local development through to production.
  • Work with Linux systems, networking, host configuration and resource contention.
  • Contribute to infrastructure security, access controls, secrets management and system hardening.

Skills

Site reliability engineering
Platform engineering
DevOps
Infrastructure engineering
Cloud environments
Linux systems
Networking fundamentals
Monitoring & observability
CI/CD
Automation

Tools

AWS
Terraform
Ansible

Job description

We're partnered with a technology company building high-performance infrastructure for decentralised financial markets.

They're looking for a Senior Site Reliability Engineer to improve the reliability, observability and operational tooling behind a latency-sensitive production platform.

The role covers production infrastructure, monitoring and alerting, incident diagnosis, deployment workflows, infrastructure automation and developer tooling. There is also a strong Linux and systems element, particularly around networking, host performance and running high-performance services in production.

The platform is still relatively early, so there is plenty of scope to improve how things are operated, introduce better automation and help set the standards the wider engineering team works to.

Responsibilities
  • Improve the reliability and operability of production systems.
  • Build and improve monitoring, logging, tracing, dashboards and alerting.
  • Improve incident diagnosis, root cause analysis and operational workflows.
  • Build safer and more repeatable deployment and rollback processes.
  • Automate repetitive operational and infrastructure work.
  • Improve CI/CD pipelines and release processes.
  • Develop internal tooling that helps engineers operate production systems more effectively.
  • Improve the developer experience from local development through to production.
  • Work with Linux systems, networking, host configuration and resource contention.
  • Contribute to infrastructure security, access controls, secrets management and system hardening.

The systems are latency-sensitive, so the role can extend into areas such as host-level tuning, kernel settings, CPU isolation and networking behaviour.

Skills & Experience
  • Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering.
  • Experience operating production infrastructure in cloud environments.
  • Strong Linux systems knowledge and understanding of networking fundamentals.
  • Experience with monitoring, observability and alerting.
  • Strong troubleshooting and root cause analysis skills.
  • Experience with CI/CD and infrastructure automation.
  • AWS, Terraform or Ansible experience would be advantageous.
  • Experience with high-performance, high-throughput or latency-sensitive systems would be particularly valuable.
  • Comfortable taking ownership of problems and driving improvements independently.
  • Significant performance-based bonus + Equity
  • Engineering-led organisation - built prioritising engineering culture
  • Direct influence over reliability, tooling and engineering practices.
  • Opportunity to work alongside a small, elite team.
  • Exposure to complex, latency-sensitive production systems.
Interested?
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Understanding Recruitment • Greater London

On-site
GBP 90,000 - 120,000
Equity
Performance-based bonus
Senior Platform Engineer
Senior Platform Engineer

Understanding Recruitment • Greater London

On-site
GBP 90,000 - 120,000
Lucrative Performance-based bonus
Equity package with significant long‑m
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

On-site
GBP 51,000 - 85,000
Bonus
Benefits
Site Reliability Engineer - Banking & Finance
Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 90,000 - 130,000
Global engineering organisation
Engineering-led culture
Technically challenging problems
+1
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000
Platform Engineer
Platform Engineer

Provn • Greater London

On-site
GBP 180,000 - 220,000
Equity
Performance bonus
Autonomy
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

On-site
GBP 65,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

The Business Connection Group • Greater London

Remote
GBP 130,000 - 150,000
Site Reliability Engineer (Tier-1 Quant Trading Organisation)
Site Reliability Engineer (Tier-1 Quant Trading Organisation)

Hamilton Barnes ? • Greater London

On-site
GBP 90,000 - 120,000