Senior Software Engineer (Application Operations)

Reward Gateway

Greater London

Hybrid

GBP 70,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Life assurance
Pension/401K
Employee Share Plan
Flexible working
Work from home bundles

Job summary

Reward Gateway is seeking a hands‑on Application Operations Engineer who bridges traditional application support and software engineering. You will execute scripted remediation, guardrail‑driven code fixes, runbook automation, and feature flag operations across AWS EKS, PHP services, and MySQL backends.

The role emphasizes shift‑left problem solving, incident responsiveness, and tooling‑driven improvements.

Qualifications

  • Hands-on capability in at least one backend language (PHP preferred; Python or similar also valuable).
  • Proven experience in application support or operations engineering in cloud environments (Kubernetes/EKS).
  • Strong experience with Datadog, Kibana, and Heap for monitoring, logging and impact analysis.

Responsibilities

  • Provide high-quality L2.5 support for PHP apps on EKS with MySQL backends within guardrails.
  • Reduce L3 escalations by automating remediation and runbook execution.
  • Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns.

Skills

Backend scripting
Kubernetes operations
Datadog
Kibana
Heap impact analysis
AWS basics
Runbooks automation

Tools

Datadog
Kibana
Heap
Jira Service Management
AWS CloudWatch

Job description

  • This hands-on role sits at the intersection of operational excellence and engineering craft
  • You’ll bridge the gap between traditional application support and software engineering by executing scripted remediation, configuration management, feature flag operations, safe, bounded code-level fixes, and runbook automation — all under clearly defined guardrails
  • The goal is to reduce unnecessary L3 escalations while increasing autonomy, quality, and impact for our Application Operations function
  • You’ll apply these practices across our AWS environment (EKS), PHP services, and MySQL databases, using Datadog as our observability platform, Kibana for log exploration, and Heap to help quantify and understand customer impact
  • Reporting To: Regional Head of Application Operations
  • Provide high-quality, timely L2.5 support for PHP applications running on EKS with MySQL backends, operating within clear guardrails that include configuration changes, feature flag operations, scripted runbooks, and safe, bounded code-level fixes
  • Model a shift-left mindset: resolve more at L2.5, automate more, and escalate less, increasing the percentage of incidents resolved without L3 involvement and improving MTTR
  • Participate in a healthy, sustainable on‑call rotation with fair schedules, clear escalation paths, and strong post‑incident learning practices
  • Apply engineering discipline to operational work: use version control, code review, and testing standards for scripts, runbooks, and automation tooling you produce
  • Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns across workloads, services, and operational scenarios
  • Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR
  • Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app‑level alerts and dashboards to minimize noise and surface actionable signals
  • Use Kibana to interrogate logs and correlate events with Datadog signals during investigations; improve log usefulness by feeding back patterns for better parsing and context
  • Use Heap to triangulate and quantify customer impact (affected flows, cohorts, and volumes) during incidents and problem investigations; incorporate findings into incident timelines and post‑incident reviews
  • Participate in service onboarding and operability reviews to ensure new and changed services meet defined supportability standards before production
  • Contribute to the Service Catalogue with accurate ownership, SLAs/SLOs, runbooks, and escalation paths for supported services
  • Act as a first responder for application incidents at L2.5: triage, diagnose, and remediate within guardrails (e.g., safe config changes, feature flag toggles, rolling restarts, cache purges, scripted data fixes). Support major incidents by providing technical context, structured diagnostics, Datadog/Kibana evidence, Heap impact analysis, and coordinated remediation alongside the incident commander
  • Use structured diagnostics before escalating — attach clear evidence, reproducibility steps, and impact assessments to every L3/SRE handoff
  • Feed operational findings into Problem Management and contribute to post‑incident reviews; capture learning in improved runbooks, alerts, and automation
  • Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first‑contact resolution, and SLO adherence
  • Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile‑aligned ways of working
  • Maintain high standards of documentation — runbooks, known errors, and operational guides are accurate, accessible, and kept up to date
  • Work closely with the Director of Application Operations, Problem Manager, and PETO peers (Platform, Infrastructure, Data, SRE) to ensure a coherent, joined‑up operational approach
  • Partner with product‑aligned engineering teams to understand application architecture, service dependencies, and failure modes; encode this knowledge into operational capabilities and runbooks
  • In scope: application‑centric remediation under guardrails; automation of known issue patterns; high‑quality runbooks; structured diagnostics; service readiness/documentation for PHP services on EKS with MySQL; ownership of app‑level dashboards/alerts in Datadog, investigative use of Kibana logs, and customer‑impact analysis via Heap
  • Standard hours are 9am - 6pm, Mon - Fri
  • 1 day in every 4 is on call, paid at 1.5x hourly rate
  • On call hours are 6pm - 9am
  • If your on call day falls on a weekend, 24 hour on call cover is required
  • What Success Looks Like
  • Increased percentage of tickets resolved at L2/L2.5, with reduced unnecessary L3 escalations
  • A maintained and actively used library of runbooks and automation scripts covering EKS /PHP / MySQL operational scenarios
  • Measurable reduction in MTTR driven by improved tooling, documentation, and automation
  • Earned trust of engineering and product peers as a technically credible, collaborative operations engineer
Benefits
  • Life assurance
  • Debt support programme & salary advances
  • Pension/401K
  • Bonus for referring a friend who we hire and have completed 3 months’ service
  • Employee Share Plan
  • Discounts to share with friends and family
  • Up to 6 months unpaid leave after five years’ service
  • Volunteer days plus a day of leave to Speak Up and be the change you want to see in the world
  • Gender neutral parental leave for primary and secondary carers
  • Unlimited free books for your professional development and one fiction book per month to help you unwind
  • Family support - Baby Bonus, caregiver support, domestic violence protection programme, parent support loan, miscarriage and baby loss support, wedding bonus and up to 3 months paid leave to take care of your family
  • Unlimited time off to give blood
  • Employee assistance programme
  • Flexible working
  • Work from home bundles
  • Trans & gender affirmation support
  • Health benefits such as free eye tests, free flu jab, freedom from addiction, menopause support, stop smoking assistance programme and health cash plan
  • Personal wellbeing allowance, personal wellbeing coach, run club, spa and gym discounts, plus cycle to work scheme
  • Bring your dog to work, drinks and breakfast

Hands‑on capability in at least one backend language (PHP preferred; Python or similar also valuable) sufficient to read, diagnose, and write safe operational scripts and minor fixes under guardrails
Proven experience in application support or operations engineering in cloud environments, ideally supporting PHP services running on Kubernetes (EKS) with MySQL backends
Practical Kubernetes skills for operations: kubectl/Helm basics, investigating pods/deployments, reading logs/events, understanding readiness/liveness probes, and performing safe rollouts/rollbacks within documented guardrails
Strong experience using Datadog (APM/metrics/traces/dashboards/alerts) for investigation and detection; confident using Kibana for log exploration and correlation; ability to leverage Heap to assess user impact and prioritize remediation
Strong communication skills; clear, concise documentation; collaborative approach focused on reducing toil, increasing automation, and raising the quality bar
MySQL operational fluency: connection and pool issues, slow query detection, query plan basics, common remediation patterns (e.g., indexing recommendations to hand to L3, safe data fixes under runbook guardrails), and understanding of replication/backup implications
Familiarity with ITSM tooling (e.g., Jira Service Management) and ITIL‑aligned incident and problem management processes
Experience with feature flag platforms and configuration‑as‑code within safe operational guardrails
Familiarity with AWS services that commonly interface with PHP/EKS workloads (e.g., CloudWatch, ALB, S3, SQS) and how they surface in Datadog and Kibana
Python
Exposure to service onboarding/operability reviews, SLOs, and contributing to a Service Catalogue
Experience balancing incident response with proactive improvement work in Agile contexts; strong documentation discipline

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer (Application Operations)
Senior Software Engineer (Application Operations)

Rewardgateway • Greater London

Hybrid
GBP 70,000 - 100,000
Hybrid work model
On-call rotation compensation
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 70,000 - 110,000
Life Insurance 4x salary
Private Medical Insurance
Employee Assistance Programme
+2
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 85,000 - 110,000
Life Insurance - 4 x Annual Salary
Private Medical Insurance
Bonus Scheme
+3
DevOps Engineer
DevOps Engineer

IT Global Consulting Ltd. • Slough

On-site
GBP 28,000 - 30,000
Operations Engineer
Operations Engineer

GradBay • Wallingford

On-site
GBP 45,000 - 65,000
Contract Tech Ops Lead
Contract Tech Ops Lead

AND Digital • City of Edinburgh

On-site
GBP 90,000 - 130,000
Lead DevOps Engineer
Lead DevOps Engineer

Priority Pass • Greater London

On-site
GBP 120,000 - 170,000
Full Stack Engineer, Platform Reliability
Full Stack Engineer, Platform Reliability

Worky • East Midlands

On-site
GBP 65,000 - 95,000
Security Analyst
Security Analyst

Doherty Associates • City Of London

On-site
GBP 35,000 - 52,000
Performance bonus
34 days annual leave
Private medical insurance
+2
Platform Support Engineer (Eng I)
Platform Support Engineer (Eng I)

Kraken Digital Asset Exchange • Greater London, Manchester

Hybrid
GBP 60,000 - 90,000