- Feeld is building a world where everyone is more intimately connected to each other and themselves. We’re an inclusive, human‑centred product with a distributed engineering team
- We’re hiring a Staff Reliability Engineer (Full Stack) to raise the reliability and operability of our production systems across backend and mobile integration
- This role exists now to improve how we detect, respond to, and prevent production incidents, and to strengthen the engineering practices (documentation, runbooks, quality bars) that help teams move quickly without compromising stability
- This is a hands‑on individual contributor role with significant cross‑team influence: you’ll lead through technical decisions, incident leadership, and pragmatic improvements to systems and process
- Reporting line: Sits on the Platform team reporting to the Head of Platform Engineering. Works closely with engineering leadership and partners day‑to‑day primarily with backend engineers
- Scope & influence: You’ll collaborate across squads to improve production ownership, reliability, and backendmobile integration patterns
- Ways of working: Remote-first, async-friendly, high trust; you’ll be expected to communicate clearly in writing and help teams adopt consistent operational practices
- Within your first year, you will have:
- Reduced incident frequency and/or impact through concrete reliability improvements (e.g., better alerting, safer deploy patterns, guardrails, playbooks)
- Made incident response more effective (clear ownership, faster MTTR, better post-incident follow‑through)
- Delivered improvements to backend/mobile integration that reduce breakages and production risk
- Established (or materially improved) documentation and operational standards that other engineers consistently use
- Own reliability outcomes across critical backend services and their integration with mobile clients (React Native)
- Lead technical problem-solving during incidents: coordinate response, diagnose root causes, communicate status, and drive to resolution
- Build and evolve monitoring/observability (dashboards, alerts, tracing, logging) that enables fast detection and diagnosis
- Drive post‑incident reviews (blameless) and ensure learnings become durable fixes (tech changes, runbooks, automation, process updates)
- Improve engineering safety and quality: guardrails, safer migrations, feature-flag practices, rollout strategies, and resilience patterns
- Partner with product, design, QA, and engineering early to align delivery plans with operational risk and reliability needs
- Strengthen documentation and onboarding: architecture notes, runbooks, service ownership docs, and “how we work” guides
- Mentor engineers through pairing, reviews, incident shadowing, and pragmatic coaching on production ownership
Benefits
- Flexible working hours
- Unlimited Paid Time Off (PTO)
- Role development & Curiosity Wallet initiatives
- GBP 3k equipment and home office budget for all new Feelders
- While we are a remote-first organisation, in-person meet ups have always been a cornerstone of the employee experience at Feeld. We provide opportunities for IRL meetups at least once a year, whether through Circle offsites or conferences, role development, coworking or team building
Demonstrated Staff-level IC leadership: influence through design reviews, technical direction, documentation, and cross-team alignmentStrong TypeScript/Node.js (or equivalent) backend experience; comfort working across services and APIsSolid observability skills: practical experience with logging/metrics/tracing and turning signals into actionable alerts and dashboardsProven incident response leadership: on‑call participation, triage, mitigation, and root‑cause analysis (RCA) with follow‑throughSignificant experience building and operating production backend systems at scale, including debugging distributed systems and performance issuesExperience collaborating with mobile teams and understanding mobilebackend integration concerns (e.g., API compatibility, releases, feature flags)React Native experience and/or strong understanding of mobile architecture patterns and release constraintsAWS (or similar cloud) experience and familiarity with infrastructure‑as‑code, CI/CD, and production toolingExperience with PostgreSQL / Redis and performance tuning in high‑traffic systemsExperience designing reliability programs (SLOs, error budgets, incident process) and running operational excellence improvementsExperience in a high-growth environment where prioritization and pragmatic trade‑offs are essential