Site Reliability Engineer

intro

Greater London

On-site

GBP 55,000 - 85,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary bonus
Compensated on-call rota
Benefits

Job summary

Intro, a fast-growing FinTech, is hiring a Site Reliability Engineer / Production Support Engineer in London. This hands-on role requires ownership of incidents, restoration of services and cross-team collaboration across Support, SRE, DevOps and Software Engineering.

You will drive monitoring and observability improvements, work with AWS or GCP environments, and read/modify backend code as needed. The role includes on-call duties and 4 days onsite commitment in London.

Qualifications

  • Strong experience in Production Support, Application Support or Production Engineering.
  • Experience at L2/L3 level investigating complex issues.
  • Knowledge of incident management, RCA, change management and production processes.
  • Hands-on experience with AWS or GCP.
  • Exposure to SRE/DevOps practices, CI/CD and modern production environments.
  • Experience with monitoring/observability tooling.
  • Comfort reading, debugging and modifying backend code.

Responsibilities

  • Take ownership of production incidents end-to-end from investigation to resolution.
  • Troubleshoot issues across applications, infrastructure, databases, APIs and distributed systems.
  • Perform RCA and help implement permanent fixes and improvements.
  • Work within ITIL/change management and incident/problem processes.
  • Improve monitoring, alerting and observability to prevent issues.
  • Collaborate with SRE/DevOps, Platform and Software Engineering teams.
  • Participate in compensated out-of-hours on-call rota.

Skills

Production Support
SRE
DevOps
Software Engineering
Incident management
L2/L3 support
AWS
GCP
CI/CD
Kubernetes
Docker
Infrastructure as Code
Monitoring
Observability
Backend code

Tools

Grafana
Prometheus
Datadog
Splunk
New Relic

Job description

Site Reliability Engineer/Production Support Engineer | FinTech |

London | 4 Days Onsite

Are you the engineer people turn to when production goes wrong?

We're partnering with a fast-growing FinTech building mission-critical banking technology for regulated financial institutions. As the platform continues to scale, they're looking for a Site reliability engineer/Production Support Engineer who can take real ownership of production while working across the boundaries of Support, SRE, DevOps and Software Engineering.

This is a hands-on production role at its core. You'll be responsible for investigating incidents, restoring services and identifying root causes, but you'll also have the opportunity to go deeper, whether that's improving observability, working with cloud infrastructure, automating processes or getting into the backend code to implement smaller fixes and improvements.

Being a growing FinTech means responsibilities aren't confined to one box. They're looking for someone who enjoys wearing multiple hats and wants genuine ownership over how production is supported and improved.

What you'll be doing
  • Taking ownership of production incidents end-to-end, from initial investigation and incident response through to resolution and post-incident review.
  • Troubleshooting complex issues across applications, infrastructure, databases, APIs and distributed systems.
  • Performing Root Cause Analysis (RCA) and helping implement permanent fixes and preventative improvements.
  • Working within structured production processes including ITIL, change management, CAB and incident/problem management.
  • Improving monitoring, alerting and observability to identify issues before they impact customers.
  • Working hands-on with AWS or GCP environments and collaborating closely with SRE/DevOps and Platform teams.
  • Reading and debugging backend code, making smaller bug fixes, modifications and production improvements where appropriate.
  • Automating repetitive operational processes and improving the efficiency of production support.
  • Supporting releases and deployments into production.
  • Working closely with Software Engineering to resolve larger or more complex application issues.
  • Participating in a compensated out-of-hours on-call rota.
What we're looking for
  • Strong experience within Production Support, Application Support or Production Engineering.
  • Experience operating at L2/L3 level, with the ability to investigate complex technical issues rather than simply escalating them.
  • Strong knowledge of incident management, incident response, RCA, change management and production support processes.
  • Hands-on experience with AWS or GCP.
  • Exposure to SRE/DevOps practices, CI/CD and modern production environments.
  • Experience with monitoring and observability tooling such as Grafana, Prometheus, Datadog, Splunk, New Relic or similar.
  • Comfortable reading, debugging and modifying backend code. The specific programming language isn't important, Python, Java, Go, C#, Kotlin or similar are all relevant.
  • Experience supporting distributed, event-driven or highly available systems.
  • Comfortable working closely with Software Engineering, DevOps/SRE and Infrastructure teams.
Nice to have
  • FinTech, banking, payments, trading or wider financial services experience.
  • Terraform or other Infrastructure as Code tooling.
  • Kubernetes and Docker.
  • Experience within a startup or scale-up environment.
  • Experience supporting high-volume or business-critical transactional platforms.

London – 4 days onsite

£55,000–£85,000 + discretionary bonus + benefits + compensated on-call

If you're a Production Support Engineer who enjoys going beyond the ticket queue, getting into the infrastructure, understanding the code and improving how production operates, I'd be keen to speak.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support Engineer
Production Support Engineer

intro • Greater London

On-site
GBP 70,000 - 95,000
On-call allowance
London onsite
FinTech Site Reliability Engineer — Production Ownership
FinTech Site Reliability Engineer — Production Ownership

intro • Greater London

On-site
GBP 55,000 - 85,000
Discretionary bonus
Compensated on-call rota
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Understanding Recruitment • Greater London

On-site
GBP 150,000 - 200,000
UK visa sponsorship
Equity
Private healthcare
Site Reliability Engineer - Banking & Finance
Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 90,000 - 130,000
Global engineering organisation
Engineering-led culture
Technically challenging problems
+1
Site Reliability Engineer
Site Reliability Engineer

Computappoint • City Of London

Hybrid
GBP 56,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

DNS INFO LTD • City Of London

On-site
GBP 70,000 - 95,000
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 90,000
Site Reliability Engineer (SRE) / Platform Engineer
Site Reliability Engineer (SRE) / Platform Engineer

Adecco • City Of London

Hybrid
Hybrid work arrangement
Competitive day rate
London-based contract
DevOps/Site Reliability Engineer - Up to £150k + Industry Leading Bonus - Elite FinTech Firm
DevOps/Site Reliability Engineer - Up to £150k + Industry Leading Bonus - Elite FinTech Firm

Hunter Bond • Greater London

Hybrid
GBP 120,000 - 150,000
Flexible hours
Hybrid working
Investment in cutting-edge tech
Site Reliability Engineer
Site Reliability Engineer

Wedo Technology Solutions Ltd. • Greater London

Remote
GBP 63,000 - 75,000