Site Reliability Engineer / Production Support

慨正橡扯

Greater London

Hybrid

GBP 60,000 - 80,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

慨正橡扯 is looking for a Site Reliability Engineer to manage incident response and oversee the offshore Production Support team in London. This hybrid position will require you to ensure system reliability and maintain high observability standards.

Ideal candidates will bring significant experience in SRE, embracing automation tools to enhance performance. You will be instrumental in protecting client interests and contributing to the company during a vital stage of growth.

Qualifications

  • Strong understanding of incident management in production environments.
  • Experience with system observability tools for debugging.
  • Ability to build automation that enhances system reliability.

Responsibilities

  • Oversee offshore Production Support team during incidents.
  • Run incident response for quick service restoration.
  • Maintain system observability standards and alert quality.

Skills

SRE or production support experience
Understanding of observability tools
Incident response accountability
Experience with offshore support teams
Hands-on reliability engineering
Active use of AI tools
Financial services experience

Job description

Site Reliability Engineer/ Production Support

Location London (Oxford Circus) | Hybrid: 2 days per week | Reports to Head of Cloud Operations

ABOUT MONUMENT

We're building something genuinely rare: a financial brand designed for the mass affluent, the professionals, entrepreneurs and ambitious savers that traditional banks have systematically underserved for decades.

We exist to make managing wealth simpler, smarter and more human, treating every client's wealth with the same care as if it were our own.

We hold over £7 billion in client savings, serve more than 100,000 clients, and were named the UK's fastest growing fintech in 2025. The momentum is real.

THE OPPORTUNITY

Monument’s production environment is the heartbeat of a licensed bank, and the SRE role is the single point of ownership when incidents occur. You will directly oversee the offshore Production Support team, run on-call and incident response, and ensure fast detection, triage and restoration of services.

This is not a passive monitoring role. You are expected to understand at a working level all of Monument’s key system flows from the services, partners and teams in play and to actively debug incidents, escalate effectively and drive permanent fixes. You will also be a builder: using AI tools for automated alert correlation, root cause analysis and runbook generation.

For the right person, this is a rare opportunity to own production reliability at a pre-IPO challenger bank, operating at the intersection of deep engineering and real commercial consequence.

WHAT YOU'LL DO
  • Directly oversee the offshore Production Support team and be the single point person when incidents occur, escalating only to Head Of when required.
  • Run on-call and incident response; ensure fast detection, triage, and restoration.
  • Maintain observability standards (logs, metrics, traces) and alert quality (low noise, high signal).
  • Understand at a working level all key system flows, the services, partners, and teams in play, and how to actively debug an incident.
  • Lead reliability engineering: resilience patterns, performance tuning, capacity planning.
  • Facilitate post-incident reviews and track actions to completion.
  • Use AI tools for automated alert correlation, root cause analysis, and runbook generation.
  • Hunt for routine/common tasks and formulate plans on how to automate and then execute them.
THE MINDSET
  • Ownership under pressure - when things go wrong, you are calm, decisive, and effective. You own the incident until it's resolved.
  • Builder - you don't just respond to incidents; you build the automation that prevents them or resolves them faster next time.
  • Deeply curious - you understand the full system landscape and how services interact. You can debug across layers.
  • Automation-first - every manual task is a candidate for automation. You actively hunt for toil and eliminate it.
  • Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source.
WHAT YOU BRING
  • Strong SRE or production support experience with accountability for incident response in a production environment.
  • Deep understanding of observability tools, alerting, logging, and distributed systems debugging.
  • Experience managing and working with offshore support teams.
  • Hands‑on experience with reliability engineering: resilience patterns, performance tuning, capacity planning.
  • Active use of AI tools for incident triage, automation, and runbook generation.
  • Ability to understand complex system flows across multiple services and third‑party integrations.
  • Experience in financial services or similarly regulated environments is a strong advantage.
WHAT'S IN IT FOR YOU

Be the person who keeps Monument running - your work directly protects clients and the business. Build automation that genuinely matters: every runbook you automate and every alert you tune makes the system more resilient. Work with modern AI tools as a core part of your daily workflow, not as a novelty. Own production reliability at a critical stage of Monument’s growth, with real responsibility and real impact.

OUR VALUES

At Monument, our values shape how we make decisions, how we treat each other when things get hard, and how we show up for clients who expect more than standard banking.

We set ambitious goals and hold ourselves to them, not because it looks good, but because our clients' outcomes depend on it. When something isn't working, we say so early, learn from it, and move. We don't wait for perfect conditions, and we don't protect egos over progress.

We work as a genuine team, which means real collaboration, honest conversations when we disagree, and shared accountability when things go wrong. We know better decisions come from different perspectives, so we actively value the range of experiences and backgrounds our people bring.

We're always asking whether there's a smarter way to do what we do, not for the sake of change, but because standing still isn't an option in the market we're in.

If that sounds like how you like to work, we'd like to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead - Savings
Tech Lead - Savings

慨正橡扯 • Greater London

Hybrid
GBP 80,000 - 100,000
28 days annual leave plus 8 bank holidays
Performance related bonus
Equity after probation
+5
Lifecycle & Content Lead: Growth & Retention
Lifecycle & Content Lead: Growth & Retention

慨正橡扯 • Greater London

On-site
Lifecycle & Content Lead
Lifecycle & Content Lead

Hackajob Ltd • Greater London

Hybrid
GBP 90,000 - 130,000
Annual leave 28d
Bonus structure
Equity after probation
+7
Lifecycle & Content Lead
Lifecycle & Content Lead

慨正橡扯 • Greater London

On-site
Lifecycle & Content Lead
Lifecycle & Content Lead

Hackajob Ltd • Slough

Hybrid
GBP 85,000 - 110,000
28 days annual leave + 8 bank holidays
Performance related bonus structure
Equity after probation
+7
Lifecycle & Content Lead
Lifecycle & Content Lead

Limelight Health • Greater London

Hybrid
GBP 70,000 - 90,000
28 days annual leave plus 8 bank holidays
Performance-related bonus structure
Equity after probation
+6
Senior Product Designer
Senior Product Designer

慨正橡扯 • Greater London

On-site
Senior Product Designer - Scale Savings at Fintech (Hybrid)
Senior Product Designer - Scale Savings at Fintech (Hybrid)

慨正橡扯 • Greater London

On-site
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

Hybrid
GBP 65,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LLOYDS BANKING GROUP • City of Edinburgh

On-site
GBP 90,000 - 120,000
Pension contribution up to 15%
Annual performance-related bonus
Share schemes including free shares
+3