Senior Software Engineer (Application Operations)

Rewardgateway

Greater London

On-site

GBP 70,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
On-call rotation compensation

Job summary

Reward Gateway, part of Edenred, is seeking a hands-on Operations Delivery engineer at L2.5 to bridge application support and software engineering in our London hybrid office. You’ll manage scripted remediation, runbooks, and feature flag operations across PHP services on EKS with MySQL, using Datadog, Kibana and Heap to measure impact.

Standard hours are 9am–6pm, with a rotating on-call schedule and 1.5x hourly pay for on-call work, ensuring resilient service for customers while reducing L3

Qualifications

  • Proven experience in application support or operations engineering in cloud environments.
  • Hands-on capability in at least one backend language (PHP preferred) to read, diagnose, and write safe operational scripts.
  • Kubernetes operations experience with EKS, including monitoring, rollouts, and safe changes.
  • MySQL back-end expertise, including connectivity, queries, and replication basics.

Responsibilities

  • Provide L2.5 support for PHP applications on EKS with MySQL backends.
  • Automate remediation, runbooks, and guardrail-based fixes to reduce L3 escalations.
  • Develop and maintain monitoring signals, dashboards, and alerts in Datadog; use Kibana for log investigations.
  • Collaborate with across teams to improve operability, service readiness, and incident response.
  • Participate in on-call rotations with fair scheduling and structured post-incident learning.

Skills

Application support
Operations engineering
Cloud environments
Kubernetes (EKS)
MySQL backends
PHP scripting
Datadog
Kibana
Heap

Tools

Datadog
Kibana
Heap

Job description

Reward Gateway, part of Edenred, is a global leader in benefits and employee engagement. We help businesses attract, engage, and retain top talent through strategic reward, recognition, and well-being solutions.

Guided by our shared missions - ‘Making the World a Better Place to Work’ and ‘Enriching Connections, For Good’ - we’re committed to transforming workplaces and improving people’s daily lives.

Our team embodies entrepreneurial spirit, innovation, and respect. We push boundaries, speak up, and stay human, fostering a culture where imagination thrives.

This role offers a hybrid work model to be present in our London office twice a week.

This hands‑on role sits at the intersection of operational excellence and engineering craft. You’ll bridge the gap between traditional application support and software engineering by executing scripted remediation, configuration management, feature flag operations, safe, bounded code‑level fixes, and runbook automation — all under clearly defined guardrails. The goal is to reduce unnecessary L3 escalations while increasing autonomy, quality, and impact for our Application Operations function. You’ll apply these practices across our AWS environment (EKS), PHP services, and MySQL databases, using Datadog as our observability platform, Kibana for log exploration, and Heap to help quantify and understand customer impact.

Key Responsibilities

L2.5 Operations Delivery

  • Provide high-quality, timely L2.5 support for PHP applications running on EKS with MySQL backends, operating within clear guardrails that include configuration changes, feature flag operations, scripted runbooks, and safe, bounded code‑level fixes.
  • Model a shift‑left mindset: resolve more at L2.5, automate more, and escape less, increasing the percentage of incidents resolved without L3 involvement and improving MTTR.
  • Participate in a healthy, sustainable on‑call rotation with fair schedules, clear escalation paths, and strong post‑incident learning practices.

Engineering Practices Within Operations

  • Apply engineering discipline to operational work: use version control, code review, and testing standards for scripts, runbooks, and automation tooling you produce.
  • Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns across workloads, services, and operational scenarios.
  • Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR.

Observability and Service Readiness

  • Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app‑level alerts and dashboards to minimize noise and surface actionable signals.
  • Use Kibana to interrogate logs and correlate events with Datadog signals during investigations; improve log usefulness by feeding back patterns for better parsing and context.
  • Use Heap to triangulate and quantify customer impact (affected flows, cohorts, and volumes) during incidents and problem investigations; incorporate findings into incident timelines and post‑incident reviews.
  • Participate in service onboarding and operability reviews to ensure new and changed services meet defined supportability standards before production.
  • Contribute to the Service Catalogue with accurate ownership, SLAs/SLOs, runbooks, and escalation paths for supported services.

Technical Operations and Incident Participation

  • Act as a first responder for application incidents at L2.5: triage, diagnose, and remediate within guardrails (e.g., safe config changes, feature flag toggles, rolling restarts, cache purges, scripted data fixes). Support major incidents by providing technical context, structured diagnostics, Datadog/Kibana evidence, Heap impact analysis, and coordinated remediation alongside the incident commander.
  • Use structured diagnostics before escalating — attach clear evidence, reproducibility steps, and impact assessments to every L3/SRE handoff.
  • Feed operational findings into Problem Management and contribute to post‑incident reviews; capture learning in improved runbooks, alerts, and automation.

Quality, Process, and Continuous Improvement

  • Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first‑contact resolution, and SLO adherence.
  • Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile‑aligned ways of working.
  • Maintain high standards of documentation — runbooks, known errors, and operational guides are accurate, accessible, and kept up to date.

Stakeholder Collaboration

  • Work closely with the Director of Application Operations, Problem Manager, and PETO peers (Platform, Infrastructure, Data, SRE) to ensure a coherent, joined‑up operational approach.
  • Partner with product‑aligned engineering teams to understand application architecture, service dependencies, and failure modes; encode this knowledge into operational capabilities and runbooks.

Scope and Interfaces (complementary to SRE)

  • In scope: application‑centric remediation under guardrails; automation of known issue patterns; high‑quality runbooks; structured diagnostics; service readiness/documentation for PHP services on EKS with MySQL; ownership of app‑level dashboards/alerts in Datadog, investigative use of Kibana logs, and customer‑impact analysis via Heap.

Working hours and practices for the team

  • Standard hours are 9am - 6pm, Mon - Fri.
  • 1 day in every 4 is on call, paid at 1.5x hourly rate
  • On call hours are 6pm - 9am
  • If your on call day falls on a weekend, 24 hour on call cover is required.
Skills

Essential technical skills

  • Proven experience in application support or operations engineering in cloud environments, ideally supporting PHP services running on Kubernetes (EKS) with MySQL backends.
  • Hands‑on capability in at least one backend language (PHP preferred; Python or similar also valuable) sufficient to read, diagnose, and write safe operational scripts and minor fixes under guardrails.
  • Practical Kubernetes skills for operations: kubectl/Helm basics, investigating pods/deployments, reading logs/events, understanding readiness/liveness probes, and performing safe rollouts/rollbacks within documented guardrails.
  • MySQL operational fluency: connection and pool issues, slow query detection, query plan basics, common remediation patterns (e.g., indexing recommendations to hand to L3, safe data fixes under runbook guardrails), and understanding of replication/
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Application Operations Engineer – Kubernetes/EKS
Senior Application Operations Engineer – Kubernetes/EKS

Rewardgateway • Greater London

Hybrid
GBP 70,000 - 100,000
Hybrid work model
On-call rotation compensation
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

SPECTRUM IT • Greater London

Hybrid
GBP 70,000 - 100,000
Life Insurance
Private Medical Insurance
Bonus Scheme
+3
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 85,000 - 110,000
Life Insurance - 4 x Annual Salary
Private Medical Insurance
Bonus Scheme
+3
Contract Tech Ops Lead
Contract Tech Ops Lead

AND Digital • City of Edinburgh

On-site
GBP 90,000 - 130,000
Operations Engineer
Operations Engineer

AGS • United Kingdom

Hybrid
GBP 51,000 - 69,000
Operations Engineer
Operations Engineer

GradBay • Wallingford

On-site
GBP 45,000 - 65,000
Senior DevOps Engineer
Senior DevOps Engineer

Candour • Manchester

On-site
GBP 70,000 - 110,000
Fully remote
Enhanced parental benefits
Platform Support Engineer (Eng I)
Platform Support Engineer (Eng I)

Kraken Digital Asset Exchange • Greater London, Manchester

Hybrid
GBP 60,000 - 90,000
Lead Service Operations Engineer (ITSM)
Lead Service Operations Engineer (ITSM)

Toyota Connected Europe • Greater London

Hybrid
GBP 90,000 - 130,000
Deployment Engineer
Deployment Engineer

Luxoft • Greater London

On-site
GBP 90,000 - 130,000