Site Reliability Engineer

Autonomai Recruitment

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Autonomai Recruitment is seeking an SRE Engineer to drive reliability, scalability, and operational excellence across critical production systems. You will work across infrastructure, software, and platform environments with a focus on resilience, automation, and engineering quality.

Collaborating with software engineering, platform, security, and infrastructure teams, you will improve service availability, observability, and operational practices in a high-performance, low-latency environment.

Qualifications

  • Must have hands-on experience operating and troubleshooting Linux-based production environments.
  • Solid understanding of distributed systems and streaming architectures.
  • Experience with messaging platforms and modern development workflows.
  • Proficiency in at least one systems-level language (C, C++, Rust) and familiarity with large-scale data systems.

Responsibilities

  • Own reliability, performance, and availability of production systems and infrastructure.
  • Support incident response, root cause analysis, and post-incident reviews.
  • Implement SRE best practices including monitoring, alerting, and service health standards.
  • Build and improve automation to reduce toil and improve deployment consistency and recovery.
  • Collaborate with engineering teams to improve system design, fault tolerance, and capacity planning.
  • Contribute to production readiness for new services and infrastructure changes.
  • Maintain observability across systems with metrics, logging, and tracing.
  • Continuously improve operational processes and platform reliability.

Skills

Linux production
Distributed systems
Messaging platforms
C/C++/Rust
Large-scale data systems
CI/CD pipelines
Containerised environments
Troubleshooting high-availability

Job description

A high-performing trading technology firm is seeking an SRE Engineer to drive reliability, scalability, and operational excellence across critical production systems. This role is suited to a hands-on engineer who operates across infrastructure, software, and platform environments, with a strong focus on resilience, automation, and engineering quality.

You will work closely with software engineering, platform, security, and infrastructure teams to improve service availability, strengthen observability, and enhance operational practices in a high-performance, low-latency environment.

Responsibilities
  • Own the reliability, performance, and availability of business-critical production systems and infrastructure
  • Support incident response, service restoration, root cause analysis, and post-incident reviews
  • Implement SRE best practices including monitoring, alerting, and service health standards
  • Build and improve automation to reduce operational toil and enhance deployment consistency and recovery
  • Partner with engineering teams to improve system design, fault tolerance, and capacity planning
  • Contribute to production readiness for new services and infrastructure changes
  • Maintain and improve observability across systems (metrics, logging, tracing)
  • Continuously improve operational processes and platform reliability
Requirements
  • Strong hands-on experience operating and troubleshooting Linux-based production environments
  • Solid understanding of distributed systems and streaming architectures
  • Experience with messaging platforms
  • Proficiency in at least one systems-level language (C, C++, Rust, or similar)
  • Familiarity with large-scale data systems
  • Experience with CI/CD pipelines, build systems, and modern development workflows
  • Working knowledge of containerised environments
  • Proven ability to troubleshoot and resolve complex production issues in high-availability systems
Preferred Profile
  • Background in a high-scale or performance-sensitive engineering environment (e.g. trading, FAANG, or similar)
  • Experience supporting low-latency or highly distributed infrastructure
  • Track record of improving automation, observability, and system reliability
  • Strong production mindset with a focus on stability, performance, and continuous improvement

This is an opportunity to work on critical, real-time systems in a high-performance environment, with direct impact on production reliability and trading outcomes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Autonomai Recruitment • England

On-site
GBP 70,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

Selby Jennings • City Of London

On-site
GBP 120,000 - 180,000
Senior Site Reliability Engineer - Real-Time Trading Systems
Senior Site Reliability Engineer - Real-Time Trading Systems

Autonomai Recruitment • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Software Engineer/ SRE (Linux)
Software Engineer/ SRE (Linux)

United States Digital Space LLC • Basingstoke

Hybrid
GBP 60,000 - 90,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Signify Technology • Greater London

Hybrid
GBP 85,000 - 120,000
Site Reliability Engineer - Banking & Finance
Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 90,000 - 130,000
Global engineering organisation
Engineering-led culture
Technically challenging problems
+1
Site Reliability Engineer
Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 80,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

DNS INFO LTD • City Of London

On-site
GBP 70,000 - 95,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1