Site Reliability Engineer (SRE)

Duffel

City Of London

On-site

GBP 50,000 - 75,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Ownership stake in the company
Focus on personal growth
Diverse workplace environment

Job summary

A dynamic technology company in travel is seeking a Site Reliability Engineer to ensure the reliability and performance of infrastructure and applications. You will collaborate closely with engineering teams to meet scalable demands and lead efforts in improving reliability monitoring. The ideal candidate has a strong background in systems engineering and is enthusiastic about both software development and infrastructure management. The position is full-time in London.

Qualifications

  • Strong experience in infrastructure and systems engineering.
  • Comfortable with debugging various issues.
  • Enthusiasm for software development and systems engineering.
  • High bar for code quality and readability.
  • Good understanding of observability and reliability practices.

Responsibilities

  • Ensure the reliability and performance of infrastructure and applications.
  • Collaborate closely with engineering teams.
  • Guide the shift towards improved reliability monitoring.

Skills

Infrastructure and systems engineering
Debugging
Software development enthusiasm
Code quality
Observability and reliability practices
Incident response
Big picture thinking
Communication skills
Collaboration

Tools

Google Cloud Platform
Terraform
Kubernetes
Grafana
Prometheus
OpenTelemetry
Honeycomb

Job description

Overview

Create the future of travel with us

Whether it’s to visit the people closest to us, starting an exciting adventure, or a career-defining business trip, travel is an essential part of our lives. Yet we've all experienced the aches and pains of getting to our destination. Today, more than 4 billion airline passengers rely on technology that hasn't kept up with the expectations of the modern connected traveller. That’s why we’ve started to rebuild the infrastructure that underpins the travel industry. We’re on a mission to unravel travel — simplifying systems and building the tools that will make the future of travel effortless.

Engineering at Duffel

We're building tools to simplify travel distribution, search and booking. What does this actually mean? It's one common and seamless API. This brings huge technical challenges as we need to design and build a beautiful API before integrating to hundreds of airlines. Along with that we need to navigate through the differing needs and systems of each airline whilst building a fantastic developer experience to go with it.

The tools used on the team include Elixir, Phoenix, Kubernetes and Google Cloud Platform.

Site Reliability Engineering at Duffel

As an SRE at Duffel, you’ll be part of a small team within engineering that is responsible for the reliability, performance, and resilience of our infrastructure and applications. You will be working closely with engineering teams to understand their needs and help meet the demands of our product as we scale globally.

What We're Looking For
  • An infrastructure and systems engineering generalist who is comfortable diving deep into the weeds on different issues. Some recent examples include: A configuration issue between Google’s Load Balancer and the HTTP server in our main Elixir application causing HTTP 5XX responses to be returned to our customers.
  • Debugging an issue in our OpenTelemetry pipelines causing us to silently drop spans.
  • An enthusiasm for both software development and systems engineering.
  • A high bar for code and configuration quality and readability.
  • A good understanding of current observability and reliability practices.
  • Experienced and comfortable in running incident response.
  • Big picture thinking - you can make trade offs on technical work streams against business impact.
  • Fantastic communication skills. You're able to articulate what you're working on and why to the team in a clear and structured way.
  • You thrive in a collaborative environment. You believe in your own methods but keep an open mind, taking suggestions and feedback onboard as well.
Technologies
  • We run our infrastructure on Google Cloud Platform, so you’ll be helping to run a few of their products such as GKE, CloudSQL for PostgreSQL, BigQuery, Memorystore (Redis) and more.
  • We manage the infrastructure and security for a segregated PCI Cardholder Data Environment, entirely managed with Google Cloud Platform services and tooling.
  • We follow an Infrastructure as Code approach to managing our infrastructure, using Terraform.
  • We follow a GitOps approach to managing our Kubernetes configuration, using ArgoCD and Helm.
  • We manage a high-availability metrics collection system using Grafana, Thanos & Prometheus. We’re in the process of transitioning to OpenTelemetry and Honeycomb for our application telemetry (traces and metrics).
  • We manage a data pipeline using Pub/Sub, Airbyte, and dbt.
Our Current Focus

We’re currently driving a big shift in how we think about and monitor reliability across the engineering organisation, with a focus on early detection of customer-impacting issues. We’re extending and standardising our use of OpenTelemetry, and introducing Honeycomb as the single place for engineers to understand how our applications are operating in production. This project involves both technical work, on the application libraries and infrastructure that make up the OpenTelemetry pipeline, and an education piece, working to change perceptions and behaviours across engineering.

The Future
  • We currently run all our services from a single European region in Google Cloud. In the medium term, for performance, reliability, and data residency reasons, we’ll be starting to think about how to (re)architect our applications and infrastructure to span multiple regions, operating globally.
  • We deploy our application multiple times a day, but deploys are all or nothing, and when we encounter issues, roll backs are slow. One way to address this would be to invest in CI/CD performance improvements, but we’d also like to explore alternative deployment strategies like Canaries, Blue/Green, and traffic mirroring, and get more comfortable testing changes in production with real customer traffic.
What you can expect from us

We're dedicated to your personal growth. Our environment is comfortable both physically, but also in that our ears are always open to any ideas, concerns and questions. We believe that everyone should have pride in their work, taking full ownership of it and its impact. That's why everyone who joins Duffel owns a share of the company.

We are an equal opportunities employer. We believe that the key to our success is employing a diverse team, that's why recruitment decisions are only based on your experience and skills. We value your ability to problem solve and build amazing things so we welcome applications for everyone – regardless of age, sex, disability, sexual orientation, race, religion or belief.

Note to recruitment agencies

Duffel does not accept speculative CV's from external parties. Any unsolicited CV's sent to us will be treated as property of Duffel, and any attached terms and conditions associated with these CV's will be null and void.

Seniority level

Not Applicable

Employment type

Full-time

Job function

Engineering and Information Technology

Industries

Transportation, Logistics, Supply Chain and Storage

London, England, United Kingdom

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Backend
Software Engineer, Backend

Duffel • City Of London

On-site
GBP 50,000 - 70,000
Personal growth opportunities
Ownership of shares in the company
Senior Customer Success Engineer - API
Senior Customer Success Engineer - API

Duffel • Greater London

On-site
GBP 50,000 - 70,000
Work From Anywhere
Duffel Travel Allowance
Generous Parental Leave
+2
Head of Technical Support
Head of Technical Support

Duffel • Greater London

Hybrid
GBP 90,000 - 130,000
Duffel Travel Allowance
Work From Anywhere (60 days)
Sabbaticals
+5
Software Engineer (Backend)
Software Engineer (Backend)

Duffel • Greater London

Hybrid
GBP 60,000 - 85,000
Competitive salary + equity
Work from anywhere policy
Duffel travel allowance
+4
Senior Engineer
Senior Engineer

United States Digital Space LLC • Greater London

Hybrid
GBP 90,000 - 150,000
Equity
Private Healthcare
Travel allowance
+2
Senior Engineer
Senior Engineer

Duffel • Greater London

Hybrid
GBP 110,000 - 170,000
Equity
Work from anywhere
Private Healthcare
+4
Software Engineer
Software Engineer

Duffel • London

On-site
GBP 60,000 - 80,000
Ownership of company shares
Personal growth focus
Equal opportunities employer
Software Engineer, Backend - Duffel, London
Software Engineer, Backend - Duffel, London

Vacations Worldwide • Greater London

Hybrid
GBP 70,000 - 110,000
Travel Support Engineer - London
Travel Support Engineer - London

Duffel • Greater London

On-site
GBP 85,000 - 115,000
Visa sponsorship available
Work From Anywhere
Duffel Travel Allowance
+1
Agency Operations Consultant
Agency Operations Consultant

Duffel • Greater London

Hybrid
GBP 40,000 - 65,000
Travel Allowance
Work From Anywhere
Sabbatical Policy
+3