Site Reliability Engineer

Autonomai Recruitment

England

On-site

GBP 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading recruitment firm is looking for a Mid-Senior Lead Site Reliability Engineer to join their core engineering group in London. The role involves designing and scaling Linux platforms for AI-driven trading, ensuring ultra-reliable and high-performance systems. Responsibilities include optimizing performance, managing Kubernetes clusters, and driving incident responses. Candidates should come from top-tier tech environments with skills in performance tuning and system automation. This is a full-time position offering a dynamic and collaborative team culture.

Qualifications

  • Experience in Linux systems engineering and site reliability engineering.
  • Skills in performance tuning and low-latency system optimization.
  • Familiarity with containerized environments and automation tools.

Responsibilities

  • Lead SRE practices for Linux platforms powering trading workloads.
  • Optimize Linux for performance and resilience.
  • Drive incident response and reliability improvement.

Skills

SRE practices
Linux performance optimization
Kubernetes management
Python for automation
Incident response

Tools

Kubernetes
Linux
HPC

Job description

Autonomai Recruitment provided pay range

This range is provided by Autonomai Recruitment. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.

Base pay range

Direct message the job poster from Autonomai Recruitment.

Connecting Ambitious Engineers to FinTech Firms Globally 🌏 🚀 💻

Role: SRE

Location: London

The ideal candidate comes from a top-tier tech environment (FAANG, elite trading, hyperscale infra). They have experience building technology 0→1, owning systems end-to-end, and working close to the metal. They will operate across everything from bare-metal Linux to modern build and observability stacks.

Overview

Join a core engineering group as Lead Site Reliability Engineer, designing and scaling Linux platforms that underpin ML/AI-driven trading. You will architect and own reliability for massive simulation, HPC, and production workloads—ensuring ultra-reliable, ultra-fast trading systems. This is a hands‑on, leadership role focused equally on technical depth, strategic decision‑making, and driving platform SRE excellence.

Key Responsibilities
  • Lead SRE practices for Linux platforms powering low-latency, high-throughput trading workloads.
  • Architect, optimize, and tune Linux for performance, resilience, and minimal latency.
  • Drive incident response, root cause analysis, and continuous reliability improvement across production systems.
  • Oversee system automation and reproducibility—build, deploy, and fleet‑manage bare‑metal Linux and containerized stacks.
  • Manage and enhance Kubernetes clusters, network configuration, and large-scale orchestration.
  • Set observability standards; expand monitoring, alerting, and performance metrics across platforms.
  • Analyze networking, kernel-level performance, and distributed systems—solving core challenges in a multi-petabyte, multi-cluster environment.
  • Build Python tools for automation, reliability engineering, and performance analysis.
  • Design highly distributed systems.
What You Will Work On
  • Ultra-reliable, high-performance trading infrastructure where every engineering optimization affects performance
  • Next-generation simulation and HPC compute pipelines, supporting ML/AI workflows at scale.
  • Integration and continuous improvement of internal and open-source tools for automation and reliability.
  • Strategic platform direction: shaping foundational systems for critical infrastructure in an elite trading environment.
Team and Culture
  • Small, autonomous Linux SRE team with direct ownership and impact.
  • Collaborative engagement with quants, researchers, and trading experts to deliver robust platforms.
  • A culture built on deep technical ownership, learning, and high standards of performance engineering.

Apply now for an informal confidential chat!

Seniority level

Mid‑Senior level

Employment type

Full‑time

Industry

Software Development

London, England, United Kingdom

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Autonomai Recruitment • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer - Banking & Finance
Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 90,000 - 130,000
Global engineering organisation
Engineering-led culture
Technically challenging problems
+1
Site Reliability Engineer (Trade Support) – Front Office Trading Technology
Site Reliability Engineer (Trade Support) – Front Office Trading Technology

Bonhill Partners • Greater London

On-site
GBP 90,000 - 150,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 80,000 - 100,000
Healthcare
Retirement Planning
Paid Volunteering Days
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Reward Gateway • Greater London

Hybrid
GBP 60,000 - 65,000
Hybrid work option
Site Reliability Engineer
Site Reliability Engineer

Computappoint • City Of London

Hybrid
GBP 56,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

Wedo Technology Solutions Ltd. • Greater London

Remote
GBP 63,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

ScaleneWorks People Solutions LLP • Bournemouth

On-site
GBP 60,000 - 80,000