Senior Site Reliability Engineer

Thought Machine

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Employee share package

Job summary

Thought Machine seeks a highly capable software engineer to join our London-based Engineering Hub. You’ll help build fault-tolerant, scalable SaaS vault products, participate in design reviews, and own parts of disaster recovery and deployment automation.

We value individuals who can drive projects to completion, work with Kubernetes, Python or Go, and collaborate across global teams while enhancing observability and cloud infrastructure on AWS or GCP.

Qualifications

  • Must have track record delivering high-impact projects with long-term scalability.
  • Strong understanding of hosting and networking design patterns.
  • Experience driving product development with minimal handholding.

Responsibilities

  • Support product engineering to build fault-tolerant, scalable applications.
  • Contribute to disaster recovery, backup, redundancy and capacity planning discussions.
  • Participate in global on-call rotation to fix bottlenecks in SaaS environments.
  • Maintain production systems hosting Vault products.
  • Contribute to feature design to improve reliability and user experience.
  • Test and update disaster recovery strategies regularly.
  • Document assets and runbooks for day-to-day operations.
  • Mentor teammates to grow technical skills and Vault knowledge.

Skills

Python
Golang
Kubernetes
AWS
GCP
Observability

Tools

Terraform
Ansible
Puppet
Docker

Job description

Thought Machine's mission is bold - to properly and permanently rid the world's banks of legacy technology. To achieve this, we have developed the foundations of modern banking through core and payments technology which run natively in the cloud. What we are attempting is hard and means we need great people working together to build great technology.

We have grown rapidly in the past few years - growing our team to more than 550 individuals across offices in London, New York, Singapore, Sydney and our newly established Engineering Hub in Lisbon. We have raised more than £500m in funding and our investors include Molten Ventures, Eurazeo, Intesa Sanpaolo, Temasek, Nyca Partners, JPMorgan Chase Strategic Investments, Standard Chartered Ventures, and more.

We have created a culture that enables our team to produce the best work in the industry while ensuring we have fun along the way. We're regularly cited as having a fantastic workplace culture and have been recognised by Sifted magazine as having one of the highest Glassdoor ratings for a UK fintech company and the industry's most generous employee share package. Named one of the world's most innovative fintechs by Global Finance Magazine, we were also recognised by the Financial Times as one of Europe's fastest-growing companies for two consecutive years-and a UK Best Employer for 2026.

Duties:
  • Supporting the product engineering teams in building highly fault-tolerant, scalable applications by participating in design discussions, engaging in RFCs and code reviews.
  • Executing various department strategies - contributing to the design and scoping work for team members around disaster recovery, backup, redundancy and capacity planning activities.
  • Being part of a global on-call rotation responsible for identifying and fixing bottlenecks in SaaS customer environments.
  • Regular maintenance of production systems that host Vault products.
  • Driving the evolution of our SaaS products by defining and designing features that foster exceptional reliability and an unparalleled user experience.
  • Implementing and regularly testing DR strategies to ensure the highest level of resilience and fault tolerance of the platform.
  • Maintain and promote high-quality written documentation of assets, processes and runbooks that are used by the team in their day-to-day operations,
  • Working with your Manager in growing team members in their technical skills as well as their understanding of Vault Products.
Requirements:
  • You have a track record of delivering high-impact projects with focus on long-term scalability, ensuring that human intervention scales sub-linearly with usage growth.
  • You possess an up-to-date understanding of design patterns relevant to hosting and networking architectures.
  • You proactively champion product development, driven by a desire to build truly exceptional products, not just solve immediate challenges.
  • You're a high-agency individual who can independently drive projects to completion by effectively scaling your individual output with the appropriate delegation of work to team members.
  • You have a strong background working in either Python, Golang having used one of these programming languages to execute a significantly sized project or initiative.
  • You have experience working with Kubernetes or other container orchestration systems.
  • You have experience with automation/configuration management, e.g. Terraform, Puppet, Chef, Ansible.
  • You have expertise in one or more of the following areas: Database Administration, Networking, Observability Tools (such as Prometheus, Jaeger) or automation infrastructure.
  • You have extensive experience working with either GCP or AWS.
Benefits:
  • Highly competitive sal
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

JobCubby • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Pension plan up to 5%
Life insurance
+6
Cloud Support Engineer
Cloud Support Engineer

Thought Machine • Greater London

On-site
GBP 60,000 - 90,000
Highly competitive salary
Pension plan (match up to 5%)
Life insurance (3x annual salary)
+10
Senior Technical Product Manager - Vault Core
Senior Technical Product Manager - Vault Core

Thought Machine • Greater London

On-site
GBP 80,000 - 100,000
Highly competitive salary and commission
Pension plan (match up to 5%)
Life insurance - three times annual salary
+8
Senior Technical Product Manager (Vault Core)
Senior Technical Product Manager (Vault Core)

Thought Machine • Greater London

On-site
GBP 90,000 - 140,000
Highly competitive salary and commission
Pension plan (match up to 5%)
Life insurance - three times annual salary
+7
Software Engineer (Infrastructure)
Software Engineer (Infrastructure)

Thought Machine • Greater London

On-site
GBP 60,000 - 80,000
Highly competitive salary and commission
Pension plan (match up to 5%)
Life insurance (3x annual salary)
+2
Senior Software Engineer (Infrastructure)
Senior Software Engineer (Infrastructure)

Thought Machine • Greater London

On-site
GBP 80,000 - 100,000
Highly competitive salary
Pension plan (match up to 5%)
Life insurance (3x annual salary)
+5
Software Engineer
Software Engineer

Thought Machine • Greater London

On-site
GBP 90,000 - 120,000
Pension plan
Life insurance
Flexible working hours
+2
Back End Engineer
Back End Engineer

Thought Machine • Greater London

On-site
GBP 60,000 - 80,000
Highly competitive salary
Pension plan (match up to 5%)
Life Insurance - 3 times annual salary
+10
Senior SRE: Cloud-Scale, Fault-Tolerant SaaS Engineer
Senior SRE: Cloud-Scale, Fault-Tolerant SaaS Engineer

Thought Machine • Greater London

On-site
GBP 90,000 - 130,000
Employee share package
Global Head of Partnerships
Global Head of Partnerships

Thought Machine • Greater London

On-site
GBP 140,000 - 180,000
Highly competitive salary
Pension plan (match up to 5%)
Life insurance - three times annual •
+8