Syfe is APAC's largest and fastest-growing digital wealth platform, trusted with over US$10 billion in assets. We are fundamentally changing how hundreds of thousands of people across Asia-Pacific build wealth through a holistic approach to managing money rather than just pushing investment products. Backed by world‑class investors and recognised as a leader in wealthtech, we are a team of passionate builders creating the future of wealth management.
About the Role
We are looking for a Senior Site Reliability Engineer to own the reliability of Syfe's production platform end‑to‑end. Syfe runs a Kubernetes‑native, multi‑region platform (Singapore, Hong Kong, Sydney) serving a regulated digital‑wealth management product, where availability, latency and trust are first‑order product features. This is a senior individual‑contributor role, not a managerial one. You will define what “reliable” means in measurable terms, build the systems and automation that keep us there, and own and lead our on‑call and incident‑response program. You will spend your time engineering reliability into the platform – through SLOs, observability and automation – rather than firefighting, and you’ll raise the bar for how the whole engineering org operates production.
What You'll Own
- Reliability targets – define and drive SLIs/SLOs and error budgets across critical services; partner with product and engineering teams to make error‑budget based decisions that balance velocity and stability.
- On‑call & incident program – own on‑call rotation, escalation policies and paging strategy; establish incident command, run blameless post‑mortems and turn RCAs into completed reliability work; drive down MTTD and MTTR.
- Platform & Kubernetes reliability – own reliability of our EKS‑based deployment platform: GitOps delivery (ArgoCD), Helm‑based release configuration and Infrastructure as Code (Terraform/OpenTofu) on AWS; make deployments safe, progressive and reversible.
- Observability – build and mature the observability stack (Datadog for production APM/RUM; Grafana, VictoriaMetrics, ClickHouse for metrics and logs); make systems debuggable with dashboards, actionable alerts and low alert noise.
- Resilience – lead capacity planning, scalability, failure‑mode analysis, disaster‑recovery and business‑continuity planning, and game‑day/chaos exercises across regions.
- Toil reduction – identify operational toil and eliminate it with automation and self‑service tooling, so reliability scales with the platform rather than with headcount.
- Production safety – strengthen rollout/rollback paths, deployment guardrails, and secrets handling (HashiCorp Vault); partner with engineering teams to harden services before they reach production.
- Engineering influence – lead by example with hands‑on engineering: design reviews, production‑readiness reviews, runbooks, documentation and mentoring – embedding SRE practices across the org.
Qualifications (Must‑have)
- Bachelor’s or Master’s degree in Computer Science, Engineering or a related field, or equivalent practical experience.
- 4–8 years in SRE, platform engineering or DevOps, with a strong senior IC track record of owning production systems.
- Production‑grade expertise with Kubernetes and containers, and a cloud platform (AWS preferred) in a distributed‑systems environment.
- Hands‑on experience defining and operating against SLIs/SLOs and error budgets, and leading incident response and blameless post‑mortems.
- Strong expertise with observability tooling – metrics, logging, tracing, dashboards and alerting (e.g., Datadog, Prometheus/Grafana or equivalents).
- Proficient in writing automation and infrastructure code (e.g., Python, Go, Shell, Terraform) and comfortable with GitOps/CI‑CD delivery.
- Solid grasp of Linux/Unix internals, networking and cloud‑native security fundamentals.
- Strong operational rigor, ownership and clear written/verbal communication, including during high‑pressure incidents.
Nice to Have
- Experience operating in regulated or high‑trust environments (fintech, payments, etc.).
- Experience running multi‑region/multi‑cluster Kubernetes at scale.
- Familiarity with ArgoCD, Helm/Helmfile, OpenTofu, Vault, Cloudflare or comparable tooling.
- Experience building self‑service developer platforms or internal reliability tooling.
- Contributions to open‑source projects or public technical content (GitHub, blogs, talks).
- Relevant cloud or Kubernetes certifications (e.g., AWS, CKA/CKS).
Come As You Are
We believe in the power of diversity and are dedicated to creating a welcoming and innovative environment for all our employees. So we embrace and encourage applications from candidates of all backgrounds and provide equal employment opportunities for all.