Site Reliability Engineer - Frontend

Capital on Tap

Greater London

On-site

GBP 70,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Private Healthcare including dental &_
Worldwide travel insurance
Anniversary Rewards (£250, £500, £750)
Salary Sacrifice Pension Scheme up to
28 days holiday
Annual Learning and Wellbeing Budget
Enhanced Parental Leave
Cycle to Work Scheme
Season Ticket Loan
6 free therapy sessions per year
Dog Friendly Offices
Free drinks and snacks

Job summary

Capital On Tap is hiring a Site Reliability Engineer to join a hybrid embedded SRE model. You will design, build, monitor and scale platforms, collaborating with XP to own architecture and paved paths for product teams.

You will manage cloud resources, implement IaC, and enhance CI/CD pipelines while improving system reliability and performance across the stack. The role emphasizes proactive incident response and strong stakeholder collaboration.

Qualifications

  • Experience managing public cloud environments and related services.
  • Proficient in writing, managing and optimising infrastructure using IaC tools like Terraform.
  • Experience with CI/CD tools, pipelines, templates and troubleshooting.
  • Proficient with containerisation (Kubernetes and Docker).
  • Experience building/deploying frontend applications with CodeMagic.
  • Knowledge or experience with Google Firebase.
  • Experience with cloud monitoring solutions.
  • Proficiency in at least one scripting language (Python, PowerShell, Go).
  • Strong communication and collaboration skills; software development background.

Responsibilities

  • Manage and automate resources in Azure, Datadog, NGINX & Cloudflare.
  • Develop, deploy and monitor Kubernetes and Serverless resources.
  • Build, manage and evolve IaC using Terraform, Helm and Go CRDs.
  • Improve systems, processes and technologies; consult stakeholders on performance.
  • Design and maintain monitoring and alerting strategies using Datadog.
  • Contribute to new application architecture and design processes.
  • Design solutions to reduce toil and automate tasks to boost productivity.
  • Create SLIs/SLOs and increase application visibility.
  • Align with Product on SLAs and core service objectives.
  • Collaborate with core teams to build reusable, automated solutions.
  • Optimise CI/CD pipelines and developer workflows with Azure DevOps, Github, Octopus Deploy, MirrorD, Flux.
  • Lead incident response and participate in post-mortems to protect customer experience.

Skills

Public cloud
Kubernetes
Docker
CI/CD
Infrastructure as Code
Frontend build tooling (CodeMagic)
Firebase
Cloud monitoring
Scripting (Python/PowerShell/Go)
Software development background
Communication

Tools

Terraform
Helm
Go CRDs
Azure DevOps
Github
Octopus Deploy
MirrorD
Flux

Job description

London, Old Street | 2 Days in Office

SRE at Capital On Tap

At Capital On Tap, we run a hybrid embedded SRE model - We aim to work closely with the teams to provide them the best support. As a Site Reliability Engineer (SRE) you will help ensure our platforms are fast, reliable, and scalable. You ’ l l design, build, and monitor systems, preventing issues before they happen.

XP, the engineering group you will be embedded into, is the platform group that builds the foundations the rest of our engineers rely on: the shared UI frameworks for web and mobile, the backend-for-frontend layer, and the internal developer platform and shared backend libraries beneath them. Our job is to make it fast and safe for product teams to ship great features, by owning the architecture and paved paths they build on top of.

What you ’ l l be doing
  • Manage and automate resources in Azure, Datadog, NGINX & Cloudflare.
  • Develop, deploy and monitor Kubernetes and Serverless resources.
  • Build, manage, and evolve IAC using Terraform, Helm and Go CRDs.
  • Improving systems, processes, and technologies; consulting stakeholders to enhance platform performance.
  • Design and maintain comprehensive monitoring and alerting strategies using Datadog.
  • Getting involved in new application architecture & design processes.
  • Designing solutions to reduce toil, automate repetitive tasks and streamline workflows to reduce manual work and boost team productivity.
  • Creating SLIs and SLOs; increasing application visibility
  • Align with the Product team on SLAs and core service objectives
  • Collaborating with core foundational teams to build and maintain reusable, automated solutions.
  • Optimising CI/CD pipelines and developer workflows using Azure DevOps, Github, Octopus Deploy, MirrorD, and Flux.
  • Leading incident response; taking an active part in communications, investigations, remediation and post-mortems to protect the customer experience.
Our Values & Culture
  • Just Pilot: We never settle for “ g ood enough ” . We pilot new ideas fast, ask questions to figure it out, and scale quickly.
  • Why Not Today? Fast is as slow as we go - speed and simplicity gives us a competitive advantage.
  • Be a Buddy: We tap in from day one to help the team, we do the right thing even if it ’ s hard.
  • Owners and Dates: We don ’ t chase people. If you own a task and agree to a date, the expectation is that it gets done.
  • Feedback: We want our employees to flourish, so we regularly provide direct and constructive feedback.
What are we looking for
  • Experience managing public cloud environments.
  • Proficient in contributing to IaC technologies involving expertise in writing, managing, and optimising infrastructure with tools such as Terraform.
  • Experience using CI/CD tools, building pipelines, templates and troubleshooting.
  • Proficient with containerisation technologies such as Kubernetes and Docker.
  • Experience building and deploying frontend applications with tools such as CodeMagic.
  • Knowledge or Experience with Google Firebase.
  • Experience working with a cloud monitoring solution.
  • Proficiency in at least one scripting language such as Python, PowerShell, Go.
  • Great communication skills with the ability to collaborate effectively.
  • Proven experience with a software development background.
Interview process
  • First stage : 30 minute intro and values call with Talent Partner
  • Second stage : 60 minute CV overview and technical chat with the SRE team lead and the SRE & Platform Engineering Manager
  • Third stage : 75 minute technical exercise & questions with the SRE lead
  • Final stage : 30 minute chat with the Engineering Manager and a lead of XP

We will try to fit in the third and final stage on the same day.

Diversity & Inclusion

We welcome, consider and encourage applications from anyone who shares our commitment to inclusivity. Join us in creating a space where authenticity thrives, and everyone can do their best work.

Great Work Deserves Great Perks

We try not to take ourselves too seriously (all the time) so we make sure our office is decked out with a pool table, arcade machine, beer tap, and a couple of office dogs thrown in for good measure.

Check out our benefits:

  • Private Healthcare including dental and opticians services through Vitality
  • Worldwide travel insurance through Vitality
  • Anniversary Rewards ( £ 2 50, £ 5 00, £ 7 50, 4-week fully paid sabbatical)
  • Salary Sacrifice Pension Scheme up to 7% match
  • 28 days holiday (plus bank holidays)
  • Annual Learning and Wellbeing Budget
  • Enhanced Parental Leave
  • Cycle to Work Scheme
  • Season Ticket Loan
  • 6 free therapy sessions per yearDog Friendly Offices
  • Free drinks and snacks in our offices
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Data
Site Reliability Engineer - Data

Capitalontap • Greater London

Hybrid
GBP 70,000 - 110,000
Private Healthcare
Worldwide travel insurance
Sabbatical rewards
+1
Site Reliability Engineer - Data
Site Reliability Engineer - Data

Capital on Tap • Greater London

Hybrid
GBP 90,000 - 120,000
Private Healthcare including dental
Worldwide travel insurance
Anniversary Rewards
+9
Cloud Engineering Team Lead
Cloud Engineering Team Lead

Capital on Tap • Greater London

On-site
GBP 90,000 - 120,000
Private healthcare including dental/ey
Annual learning budget
Dog-friendly offices
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Reward Gateway • Greater London

Hybrid
GBP 60,000 - 65,000
Hybrid work option
Senior Data Reliability Engineer
Senior Data Reliability Engineer

Elliptic Enterprises Limited • Greater London

Hybrid
GBP 90,000 - 150,000
Hybrid working option
Remote working budget
Learning & Development budget
+3
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

Hybrid
GBP 140,000 - 200,000
ESPP
Life assurance
Income protection
+11
Site Reliability Engineer SRE
Site Reliability Engineer SRE

Client Server • Cambridge

Hybrid
GBP 63,000 - 77,000
Pension
Private Medical Insurance
Life Assurance
+5
Site Reliability Engineer
Site Reliability Engineer

Wedo Technology Solutions Ltd. • Greater London

Remote
GBP 63,000 - 75,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Reward Gateway • Greater London

Hybrid
GBP 70,000 - 110,000
Life assurance
Pension
Employee Share Plan
+3