- This hands-on role sits at the intersection of operational excellence and engineering craft
- You’ll bridge the gap between traditional application support and software engineering by executing scripted remediation, configuration management, feature flag operations, safe, bounded code-level fixes, and runbook automation — all under clearly defined guardrails
- The goal is to reduce unnecessary L3 escalations while increasing autonomy, quality, and impact for our Application Operations function
- You’ll apply these practices across our AWS environment (EKS), PHP services, and MySQL databases, using Datadog as our observability platform, Kibana for log exploration, and Heap to help quantify and understand customer impact
- Reporting To: Regional Head of Application Operations
- Provide high-quality, timely L2.5 support for PHP applications running on EKS with MySQL backends, operating within clear guardrails that include configuration changes, feature flag operations, scripted runbooks, and safe, bounded code-level fixes
- Model a shift-left mindset: resolve more at L2.5, automate more, and escalate less, increasing the percentage of incidents resolved without L3 involvement and improving MTTR
- Participate in a healthy, sustainable on‑call rotation with fair schedules, clear escalation paths, and strong post‑incident learning practices
- Apply engineering discipline to operational work: use version control, code review, and testing standards for scripts, runbooks, and automation tooling you produce
- Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns across workloads, services, and operational scenarios
- Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR
- Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app‑level alerts and dashboards to minimize noise and surface actionable signals
- Use Kibana to interrogate logs and correlate events with Datadog signals during investigations; improve log usefulness by feeding back patterns for better parsing and context
- Use Heap to triangulate and quantify customer impact (affected flows, cohorts, and volumes) during incidents and problem investigations; incorporate findings into incident timelines and post‑incident reviews
- Participate in service onboarding and operability reviews to ensure new and changed services meet defined supportability standards before production
- Contribute to the Service Catalogue with accurate ownership, SLAs/SLOs, runbooks, and escalation paths for supported services
- Act as a first responder for application incidents at L2.5: triage, diagnose, and remediate within guardrails (e.g., safe config changes, feature flag toggles, rolling restarts, cache purges, scripted data fixes). Support major incidents by providing technical context, structured diagnostics, Datadog/Kibana evidence, Heap impact analysis, and coordinated remediation alongside the incident commander
- Use structured diagnostics before escalating — attach clear evidence, reproducibility steps, and impact assessments to every L3/SRE handoff
- Feed operational findings into Problem Management and contribute to post‑incident reviews; capture learning in improved runbooks, alerts, and automation
- Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first‑contact resolution, and SLO adherence
- Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile‑aligned ways of working
- Maintain high standards of documentation — runbooks, known errors, and operational guides are accurate, accessible, and kept up to date
- Work closely with the Director of Application Operations, Problem Manager, and PETO peers (Platform, Infrastructure, Data, SRE) to ensure a coherent, joined‑up operational approach
- Partner with product‑aligned engineering teams to understand application architecture, service dependencies, and failure modes; encode this knowledge into operational capabilities and runbooks
- In scope: application‑centric remediation under guardrails; automation of known issue patterns; high‑quality runbooks; structured diagnostics; service readiness/documentation for PHP services on EKS with MySQL; ownership of app‑level dashboards/alerts in Datadog, investigative use of Kibana logs, and customer‑impact analysis via Heap
- Standard hours are 9am - 6pm, Mon - Fri
- 1 day in every 4 is on call, paid at 1.5x hourly rate
- On call hours are 6pm - 9am
- If your on call day falls on a weekend, 24 hour on call cover is required
- What Success Looks Like
- Increased percentage of tickets resolved at L2/L2.5, with reduced unnecessary L3 escalations
- A maintained and actively used library of runbooks and automation scripts covering EKS /PHP / MySQL operational scenarios
- Measurable reduction in MTTR driven by improved tooling, documentation, and automation
- Earned trust of engineering and product peers as a technically credible, collaborative operations engineer
Benefits
- Life assurance
- Debt support programme & salary advances
- Pension/401K
- Bonus for referring a friend who we hire and have completed 3 months’ service
- Employee Share Plan
- Discounts to share with friends and family
- Up to 6 months unpaid leave after five years’ service
- Volunteer days plus a day of leave to Speak Up and be the change you want to see in the world
- Gender neutral parental leave for primary and secondary carers
- Unlimited free books for your professional development and one fiction book per month to help you unwind
- Family support - Baby Bonus, caregiver support, domestic violence protection programme, parent support loan, miscarriage and baby loss support, wedding bonus and up to 3 months paid leave to take care of your family
- Unlimited time off to give blood
- Employee assistance programme
- Flexible working
- Work from home bundles
- Trans & gender affirmation support
- Health benefits such as free eye tests, free flu jab, freedom from addiction, menopause support, stop smoking assistance programme and health cash plan
- Personal wellbeing allowance, personal wellbeing coach, run club, spa and gym discounts, plus cycle to work scheme
- Bring your dog to work, drinks and breakfast
Hands‑on capability in at least one backend language (PHP preferred; Python or similar also valuable) sufficient to read, diagnose, and write safe operational scripts and minor fixes under guardrails
Proven experience in application support or operations engineering in cloud environments, ideally supporting PHP services running on Kubernetes (EKS) with MySQL backends
Practical Kubernetes skills for operations: kubectl/Helm basics, investigating pods/deployments, reading logs/events, understanding readiness/liveness probes, and performing safe rollouts/rollbacks within documented guardrails
Strong experience using Datadog (APM/metrics/traces/dashboards/alerts) for investigation and detection; confident using Kibana for log exploration and correlation; ability to leverage Heap to assess user impact and prioritize remediation
Strong communication skills; clear, concise documentation; collaborative approach focused on reducing toil, increasing automation, and raising the quality bar
MySQL operational fluency: connection and pool issues, slow query detection, query plan basics, common remediation patterns (e.g., indexing recommendations to hand to L3, safe data fixes under runbook guardrails), and understanding of replication/backup implications
Familiarity with ITSM tooling (e.g., Jira Service Management) and ITIL‑aligned incident and problem management processes
Experience with feature flag platforms and configuration‑as‑code within safe operational guardrails
Familiarity with AWS services that commonly interface with PHP/EKS workloads (e.g., CloudWatch, ALB, S3, SQS) and how they surface in Datadog and Kibana
Python
Exposure to service onboarding/operability reviews, SLOs, and contributing to a Service Catalogue
Experience balancing incident response with proactive improvement work in Agile contexts; strong documentation discipline