Senior Site Reliability Engineer (APAC)

REAP Limited

Hong Kong

On-site

HKD 600,000 - 1,200,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

REAP Limited is building a Site Reliability Engineering practice to own the reliability of a payments platform across multi-region AWS. You will lead by example, drive technical excellence, and help engineers level up in a distributed, flat team.

You will set SLIs/SLOs, contribute to incident response, and push for full IaC coverage while automating provisioning and observability. This role demands hands-on expertise in AWS, Terraform, Kubernetes, and PCI DSS environments.

Qualifications

  • Real SRE practice with defined SLOs, error budgets or incident postmortems.
  • Strong Linux and system fundamentals including networking and admin tasks.
  • Deep Terraform skills: state management, modules, and large-scale IaC migration or greenfield buildouts.
  • Advanced AWS knowledge in multi-account/multi-region environments (Control Tower, IAM, VPC, RDS).
  • Experience operating ECS/Fargate and production Kubernetes at scale, including cluster lifecycle and upgrades.
  • Serverless and event-driven systems: Lambda, SQS, EventBridge, with failure mode awareness.
  • GitOps and CI/CD pipelines with practical tooling (Argo CD, GitHub Actions).
  • Observability discipline: logging, metrics, tracing with modern tools and standards adoption.
  • Languages: Python, Go or Bash for automation.
  • Ownership across a service's life cycle, including security and compliance in payments.

Responsibilities

  • Define SLIs/SLOs in collaboration with product and engineering teams.
  • Participate in on-call rotations and mature incident response and blameless postmortems.
  • Drive Infrastructure as Code coverage, reducing manual provisioning and drift.
  • Consolidate Terraform estate into modular, governed codebases with drift detection.
  • Automate account provisioning and regional standups for repeatable deployments.
  • Build self-service interfaces so product teams can provision, deploy, observe without tickets.
  • Design epoxy environments for isolated production-like testing and automatic cleanup.
  • Implement industry-standard observability and actionable alerting across platforms.
  • Own cloud operations for PCI DSS and regulated financial systems.
  • Embed security into infra: secrets, least privilege, network segmentation.
  • Support AI-assisted engineering with tooling for code, reviews, and incident analysis.
  • Collaborate with product teams to treat the platform as a product and align with customers.

Skills

SRE practice
Linux fundamentals
Terraform
AWS multi-account
ECS/Fargate & Kubernetes
Serverless & event-driven
GitOps & CI/CD
Observability
Python/Go/Bash
PCI DSS experience
Communication
AI-assisted engineering
Production incidents management

Tools

Terraform
CloudFormation
Kubernetes
ECS/Fargate
Argo CD
GitHub Actions
New Relic
CloudWatch
Datadog
Prometheus

Job description

Reap is building a Site Reliability Engineering practice, and this role is central to it.

We run card issuing, payouts, FX and stablecoin settlement across multiple AWS regions, under PCI DSS and financial regulation. The platform is growing quickly — into new markets, new products, and now agent-initiated payments — and the infrastructure underneath it needs to grow up with it. That means real service ownership, reliability measured in SLIs and SLOs rather than intuition, and a platform that engineering teams can serve themselves from instead of queueing for.

That work is largely still ahead of us, which is the appeal. You will help decide what reliability means at Reap, what the platform looks like, and what good engineering practice is in this domain — rather than inheriting someone else's answers.

This is a deeply hands‑on senior individual contributor role. We expect you to lead by example and drive technical excellence through the systems you build, the standards you set, and the way you help the engineers around you level up. The team is deliberately flat and distributed across time zones.

Technologies You’ll Use

Cloud: AWS, multi-account across multiple regions — Transit Gateway, PrivateLink, site-to-site VPN, WAF, KMS

Compute: ECS/Fargate, Lambda and EKS

Infrastructure as Code: Terraform and CloudFormation

CI/CD: GitHub Actions, Argo CD

Messaging and streaming: SQS, EventBridge, Kafka

Observability: New Relic, CloudWatch

Languages and scripting: Python, Go and Bash

What You’ll Do

As a Senior Site Reliability Engineer at Reap, you will join a team of experienced engineers transitioning a DevOps organisation into Site Reliability Engineering. You will own the reliability of a payments platform while rebuilding the foundation it runs on — the lights stay on while the platform gets replaced underneath them. We treat repetitive manual work as a bug in the platform, not a chore for a human, and your job is to delete whole categories of it rather than absorb them faster.

Define what reliability means here: set SLIs and SLOs with product and engineering teams, introduce error budgets, and make "ship or stabilise?" a matter of arithmetic rather than argument

Take part in our on-call rotation, and help build the incident response and blameless postmortem practice around it

Drive Reap to full Infrastructure as Code coverage: bring the remaining legacy infrastructure under IaC, and get to no manual provisioning, no drift, and every resource defined, versioned and reproducible

Consolidate our Terraform estate into a coherent, well‑structured codebase with module standards, governance and automated drift detection

Automate account provisioning and environment setup so that new regions and services can be stood up repeatably and consistently

Build the golden paths and self‑service interfaces that let product and engineering teams provision, deploy and observe without filing a ticket, with escape hatches for the cases they do not cover

Design and implement an epoxy environment platform so any developer, and any coding agent, can get an isolated production‑like environment on demand and have it cleaned up automatically

Implement industry-standard observability across logging, metrics and distributed tracing, with alerting that is actionable and trusted

Own cloud operations for PCI DSS and regulated financial systems — uptime, failover, capacity, disaster recovery and incident response

Embed security into the infrastructure layer: secrets management, least‑privilege IAM, network segmentation and compliance controls

Build infrastructure that lets AI assistants and agents operate safely: sane blast radius, strong isolation, auditable actions

Partner with product and engineering teams so that reliability work is negotiated rather than imposed — we treat the platform as a product and our engineering teams as its customers

Skills We’re Looking For

Real SRE practice, not just the vocabulary: you have defined SLOs, run error budgets, or built an incident and postmortem process that people actually used

Strong Linux and computer‑systems fundamentals: internals, networking, and administration. This work rests on them

Deep Terraform: state management, module design, and enforcing standards across a team. You have owned a full IaC migration or a greenfield buildout at scale

Strong AWS across multi‑account, multi‑region estates: Control Tower, IAM, networking, RDS, and cost management

Containers and orchestration: comfortable operating ECS and Fargate, and able to build and run production Kubernetes at scale — cluster lifecycle and upgrades, autoscaling, resource limits and multi‑tenant workload isolation

Serverless and event‑driven systems: Lambda, SQS and EventBridge in production, including the failure modes that only show up at scale — retries, poison messages, ordering and idempotency

GitOps and delivery: Argo CD or equivalent, and CI/CD pipelines with GitHub Actions that engineers actually enjoy using

Observability in practice: logging, metrics and tracing with tools such as New Relic, CloudWatch, Datadog or Prometheus — and getting standards adopted, not just published

Python, Go or Bash for automation and tooling

Full‑lifecycle ownership: you own a service across its whole life — from understanding the need, through design and delivery, to observability, incident response, capacity, upgrades and security patching

Experience operating in a regulated environment, and comfort with the constraints that come with PCI DSS and financial compliance

Self‑management and cross‑team influence: you can run your own projects and build alignment without authority. Nobody will sequence your work for you

Communication as a first‑class skill, weighted equally with technical depth. You can take a problem, explain your approach in plain language, and decompose it into work — for a person or for an agent

AI‑assisted engineering: you already use AI tools seriously for code, review, automation, incident analysis and documentation, and you have opinions about where they work and where they do not

Genuine comfort with production incidents. In payments, high blast‑radius incidents preempt everything

Bonus Skills

Significant experience in SRE, DevOps or infrastructure engineering, including time in an organisation with a mature SRE practice — defined SLOs, self‑service deploys, and real on‑call

Experience in fintech, payments, card issuing or another regulated environment

Hands‑on PCI DSS work: designing cardholder‑data environments, or reducing scope

Kubernetes or EKS inside a PCI DSS environment: isolating the cardholder‑data environment through namespace and network‑policy segmentation, dedicated node groups, admission control and pod security standards — and the audit logging that makes it provable

Having designed and operated ephemeral or on‑demand environment platforms, including the hard parts — database seeding, service dependencies and secrets in short‑lived environments

Having built an internal developer platform, account factory or landing zone from scratch

Configuration management with Ansible or similar

AWS certifications

Curiosity about stablecoins and the Web2 / Web3 intersection

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (APAC)
Senior Site Reliability Engineer (APAC)

Reap • Hong Kong

On-site
HKD 900,000 - 1,500,000
Senior SRE - APAC Payments Platform, Self-Service IaC
Senior SRE - APAC Payments Platform, Self-Service IaC

REAP Limited • Hong Kong

On-site
HKD 600,000 - 1,200,000
Tech Lead, Platform Team (Backend-Focused)
Tech Lead, Platform Team (Backend-Focused)

REAP Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
Insurance coverage after probation
Reap Card stipend
AI tools at work
+1
Tech Lead, Platform Team (Backend-Focused)
Tech Lead, Platform Team (Backend-Focused)

Reap • Hong Kong

Hybrid
HKD 900,000 - 1,500,000
Hybrid/remote work environment
Insurance after probation
Reap Card stipend
+1
APAC Senior SRE: Payments Platform Reliability
APAC Senior SRE: Payments Platform Reliability

Reap • Hong Kong

On-site
HKD 900,000 - 1,500,000
Engineering Manager, AI (Agentic Enablement)
Engineering Manager, AI (Agentic Enablement)

REAP Limited • Hong Kong

Hybrid
HKD 1,000,000 - 1,400,000
Insurance after probation
Reap Card stipend
AI tools at work
Senior Software Engineer, Reap Direct
Senior Software Engineer, Reap Direct

Reap • Hong Kong

On-site
HKD 600,000 - 800,000
Senior Product Manager, Agentic Enablement (Platform AI)
Senior Product Manager, Agentic Enablement (Platform AI)

REAP Limited • Hong Kong

Hybrid
HKD 900,000 - 1,500,000
Insurance coverage after probation
Reap Card stipend
Use of AI tools at work
+1
Staff Engineer, Embedded Finance & AI
Staff Engineer, Embedded Finance & AI

REAP Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
Reap Card stipend
Engineering Manager, AI (Agentic Enablement)
Engineering Manager, AI (Agentic Enablement)

Reap • Hong Kong

Hybrid
HKD 900,000 - 1,300,000
Reap Card stipend
Insurance coverage after probation
Flexible hybrid/remote work
+1