SRE: Reliability & Observability Engineer

Trainline

Greater London

Hybrid

GBP 55,000 - 63,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Private healthcare
Dental insurance
Work from abroad policy
Share purchase plans (2-for-1)
EV Scheme
Extra festive time off
Family-friendly benefits

Job summary

Trainline is seeking a mid-level Site Reliability Engineer to help drive reliability across the platform. You will join the Reliability & Operations Engineering team and contribute to incident response, post-incident reviews and on-call rotations.

You'll design, build and maintain observability using metrics, logs, events and traces, while improving monitoring and alerting aligned with customer impact. A hybrid model and strong AWS/Infra as code focus shape the role.

Qualifications

  • Experience of SRE concepts such as SLI, SLO and error budgets.
  • Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar
  • Experience working with cloud providers (preferably AWS).
  • Experience troubleshooting Linux operating systems.
  • Experience of scripting in at least one language (preferably Python)
  • Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts.
  • Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling).
  • Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions.
  • Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform.

Responsibilities

  • Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform.
  • Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration.
  • Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience.
  • Taking part in the SRE on-call rotation.
  • Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis.
  • Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD).
  • Ensuring relevant operational data is surfaced quickly and clearly during live incidents.
  • Making informed tooling and technology choices using SRE principles, balancing team and business needs.
  • Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling.
  • Collaborating with product engineering teams to ensure services are operationally ready and deployed safely.
  • Advising on reliability and resilience practices.
  • Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals.
  • Prioritising work effectively and collaborating using agile processes to deliver against team and business goals.

Skills

SRE concepts (SLI/SLO)
Observability tooling
AWS cloud experience
Linux troubleshooting
Scripting (Python)
Load balancing & reverse proxy
Application architecture concepts
Time series data & metrics
Build/deploy tooling (GitHub Actions,

Tools

New Relic
ELK Stack
Grafana
Incident.io
Docker
ECS
Terraform
GitHub Actions
AWS

Job description

Trainline is seeking a mid-level Site Reliability Engineer to help drive reliability across the platform. You will join the Reliability & Operations Engineering team and contribute to incident response, post-incident reviews and on-call rotations.

You'll design, build and maintain observability using metrics, logs, events and traces, while improving monitoring and alerting aligned with customer impact. A hybrid model and strong AWS/Infra as code focus shape the role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000
SRE Lead: Observability & Platform Reliability
SRE Lead: Observability & Platform Reliability

Camwebdir • United Kingdom

Remote
GBP 90,000 - 120,000
23 days holiday
Birthday day off
Private medical insurance
+5
Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Site Reliability Engineer
Site Reliability Engineer

Trainline • Greater London

Hybrid
GBP 55,000 - 63,000
Private healthcare
Dental insurance
Work from abroad policy
+4
SRE Manager: Scale, Reliability & Observability Leader
SRE Manager: Scale, Reliability & Observability Leader

UST • Nottingham

On-site
GBP 90,000 - 120,000
SRE Engineer – FinTech Reliability, Observability & Cloud
SRE Engineer – FinTech Reliability, Observability & Cloud

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 90,000 - 130,000
Global engineering organisation
Engineering-led culture
Technically challenging problems
+1
SRE Manager: Reliability & Incident Leadership (Hybrid London)
SRE Manager: Reliability & Incident Leadership (Hybrid London)

Gravitas Recruitment Group (Global) Ltd • Greater London

Hybrid
GBP 75,000 - 100,000
Site Reliability Engineer (DV Security Clearance)
Site Reliability Engineer (DV Security Clearance)

Onyx-Conseil • Manchester

Hybrid
GBP 90,000 - 120,000
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1