Site Reliability Engineer

Trainline

Greater London

Hybrid

GBP 55,000 - 63,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Private healthcare
Dental insurance
Work from abroad policy
Share purchase plans (2-for-1)
EV Scheme
Extra festive time off
Family-friendly benefits

Job summary

Trainline is seeking a mid-level Site Reliability Engineer to help drive reliability across the platform. You will join the Reliability & Operations Engineering team and contribute to incident response, post-incident reviews and on-call rotations.

You'll design, build and maintain observability using metrics, logs, events and traces, while improving monitoring and alerting aligned with customer impact. A hybrid model and strong AWS/Infra as code focus shape the role.

Qualifications

  • Experience of SRE concepts such as SLI, SLO and error budgets.
  • Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar
  • Experience working with cloud providers (preferably AWS).
  • Experience troubleshooting Linux operating systems.
  • Experience of scripting in at least one language (preferably Python)
  • Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts.
  • Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling).
  • Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions.
  • Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform.

Responsibilities

  • Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform.
  • Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration.
  • Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience.
  • Taking part in the SRE on-call rotation.
  • Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis.
  • Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD).
  • Ensuring relevant operational data is surfaced quickly and clearly during live incidents.
  • Making informed tooling and technology choices using SRE principles, balancing team and business needs.
  • Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling.
  • Collaborating with product engineering teams to ensure services are operationally ready and deployed safely.
  • Advising on reliability and resilience practices.
  • Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals.
  • Prioritising work effectively and collaborating using agile processes to deliver against team and business goals.

Skills

SRE concepts (SLI/SLO)
Observability tooling
AWS cloud experience
Linux troubleshooting
Scripting (Python)
Load balancing & reverse proxy
Application architecture concepts
Time series data & metrics
Build/deploy tooling (GitHub Actions,

Tools

New Relic
ELK Stack
Grafana
Incident.io
Docker
ECS
Terraform
GitHub Actions
AWS

Job description

About us

At Trainline, our purpose is to empower greener travel choices, connecting people and places. Trainline enables millions of travellers to find and book the best value tickets across carriers, fares, and journey options through our highly rated mobile app, website, and B2B partner channels.

Great journeys start with Trainline

We’re Europe’s leading independent rail platform, helping millions of travellers find and book the best-value rail and coach journeys across our app, website and partner channels.

Our job is to make the green travel choice the best choice. By building a better train travel experience, we help more people choose rail - creating a positive impact for customers, our business and the planet.

We’re a team of more than 1,000 Trainliners from over 50 nationalities, working across London, Paris, Barcelona, Milan, Edinburgh and Madrid. Now is a brilliant time to join us and help shape the future of travel.

Introducing Reliability & Operations Engineering

Trainline is a fast-growing tech company powering world-class digital journeys for millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices.

The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery, respond to incidents, and continuously strengthen system reliability.

We’re looking for a mid-level Site Reliability Engineer to help drive this forward. You’ll bring solid production experience, a growth mindset, and a willingness to challenge and be challenged — contributing to platform reliability while developing broader technical ownership with support from senior engineers.

As an SRE at Trainline, you'll be working on...
  • Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform

  • Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration

  • Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience

  • Taking part in the SRE on-call rotation

  • Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis

  • Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD)

  • Ensuring relevant operational data is surfaced quickly and clearly during live incidents

  • Making informed tooling and technology choices using SRE principles, balancing team and business needs

  • Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling

  • Collaborating with product engineering teams to ensure services are operationally ready and deployed safely

  • Advising on reliability and resilience practices

  • Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals

  • Prioritising work effectively and collaborating using agile processes to deliver against team and business goals

Our Tech Stack
  • AWS

  • New Relic

  • ELK stack

  • Grafana

  • Incident.io

  • Docker, ECS

  • Terraform

  • Github Actions

We'd love to hear from you if you have...
  • Experience of SRE concepts such as SLI, SLO and error budgets.

  • Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar

  • Experience working with cloud providers (preferably AWS).

  • Experience troubleshooting Linux operating systems.

  • Experience of scripting in at least one language (preferably Python)

  • Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts.

  • Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling).

  • Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions.

  • Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform.

More information:

Enjoy fantastic perks like private healthcare & dental insurance, a generous work from abroad policy, 2-for-1 share purchase plans, an EV Scheme to further reduce carbon emissions, extra festive time off, and excellent family-friendly benefits.

We prioritise career growth with clear career paths, transparent pay bands, personal learning budgets, and regular learning days. Jump on board and supercharge your career from day one!

We're operating a hybrid model and ask that Trainliners work from the office a minimum of 60% of their time over a 12-week period. We also have a 28-day Work from Abroad policy.

Our values represent the things that matter most to us and what we live and breathe everyday, in everything we do:

  • Think Big - We're building the future of rail

  • Own It - We focus on every customer, partner and journey

  • Travel Together - We're one team

  • Do Good - We make a positive impact

We know that having a diverse team makes us better and helps us succeed. And we mean all forms of diversity - gender, ethnicity, sexuality, disability, nationality and diversity of thought. That's why we're committed to creating inclusive places to work, where everyone belongs and differences are valued and celebrated.

Interested in finding out more about what it's like to work at Trainline? Why not check us out on LinkedIn, Instagram and Glassdoor!

Compensation: £55K – £63K

  • Base Salary £55K – £63K

Find more English Speaking Jobs in United Kingdom on Arbeitnow

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer - Platform
Engineer - Platform

Trainline • Greater London

Hybrid
GBP 60,000 - 70,000
Private healthcare
Dental insurance
EV scheme
+3
Junior .NET Engineer
Junior .NET Engineer

Trainline • Greater London

Hybrid
GBP 35,000 - 38,000
Private healthcare
Dental insurance
Work from abroad policy
+4
IT Support Analyst
IT Support Analyst

Trainline • Greater London

On-site
GBP 38,000 - 42,000
Private healthcare
Dental insurance
Work from abroad policy
+3
Junior .NET Engineer
Junior .NET Engineer

Whereby • Greater London

Hybrid
GBP 35,000 - 55,000
Private healthcare
Dental insurance
Work from abroad policy
+4
Senior Backend Engineer - .Net
Senior Backend Engineer - .Net

Trainline • Greater London

Hybrid
GBP 85,000 - 95,000
Private healthcare
Dental insurance
Work from abroad policy
+4
Service Desk Engineer
Service Desk Engineer

Trainline • City of Edinburgh

Hybrid
GBP 32,000 - 48,000
Private healthcare
Dental insurance
Work from abroad policy
+4
Junior .NET Engineer
Junior .NET Engineer

Trainline plc • Greater London

Hybrid
GBP 32,000 - 52,000
Private healthcare
Dental insurance
Work from abroad policy
+4
Senior IT Support Analyst
Senior IT Support Analyst

Trainline • City Of London

On-site
GBP 35,000 - 50,000
Private healthcare
Generous work-from-abroad policy
2-for-1 share purchase plan
Engineer - .NET Backend (London)
Engineer - .NET Backend (London)

Trainline • Greater London

On-site
GBP 55,000 - 85,000
Private healthcare
Dental insurance
Work from abroad policy
+4
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Trainline • Greater London

On-site
GBP 110,000 - 160,000
Private healthcare
Dental insurance
Work from abroad policy
+4