Senior Site Reliability Engineer

Kody

Hong Kong

On-site

HKD 900,000 - 1,500,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive Package
A dynamic and innovative team
Collaborative, inclusive working env.

Job summary

Kody is seeking a Senior Site Reliability Engineer to own the reliability and scalability of our global payments platform. Based in Hong Kong or Shenzhen, you will oversee observability, incident response, SLOs, and cloud infrastructure across Europe, Asia, and North America.

You will mentor engineers, drive automation to reduce toil, and partner with global teams to ensure security and uptime in regulated payment environments.

Qualifications

  • 8+ years of hands-on SRE/DevOps experience in high-availability production systems.
  • Strong AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, and Linux skills.
  • Experience with distributed systems, high availability, disaster recovery, capacity planning.
  • Domain experience in payments/fintech with PCI-DSS and security focus.
  • Excellent English communication; based in Hong Kong or Shenzhen.

Responsibilities

  • Incident Management & On-Call: lead incident response for Sev1/Sev2 events to minimize MTTR.
  • Production Operations: diagnose, triage, mitigate, and resolve incidents across payment services, Kubernetes, databases, messaging, and cloud.
  • SLO & Reliability Engineering: define and maintain SLOs, SLIs, error budgets, alerts and readiness processes.
  • Continuous Optimization: drive reliability improvements via automation, observability, capacity planning, and RCA.
  • Security & Compliance: partner with global teams to strengthen PCI-DSS compliant resilience.
  • Technical Leadership: mentor juniors and reduce toil through automation.

Skills

AWS
Kubernetes (EKS)
Terraform
PostgreSQL
Redis
Kafka
Linux
Datadog
Prometheus
Grafana
SRE Principles

Tools

Datadog
Prometheus
Grafana

Job description

Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Job Summary

Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Key Responsibilities
  • Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.
Requirements
  • Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Location & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.
Leadership & Operational Excellence
  • Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
  • Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
  • Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
  • Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.
Benefits
  • - Competitive Package
  • - A dynamic and innovative team
  • - Collaborative, inclusive working environment
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Global Payments Platform
Senior Site Reliability Engineer - Global Payments Platform

Kody • Hong Kong

On-site
HKD 900,000 - 1,500,000
Competitive Package
A dynamic and innovative team
Collaborative, inclusive working env.
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

TEKsystems • Hong Kong

On-site
HKD 500,000 - 900,000
Sr. Manager, Site Reliability & Innovation, IT
Sr. Manager, Site Reliability & Innovation, IT

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000
Senior Site Reliability Engineer (APAC)
Senior Site Reliability Engineer (APAC)

Reap • Hong Kong

On-site
HKD 900,000 - 1,500,000
Backend Kotlin Developer (Senior)
Backend Kotlin Developer (Senior)

Kody • Hong Kong

On-site
HKD 480,000 - 720,000
Equity available
Global tech-driven culture
Career growth opportunities
Senior DevOps Engineer- Financial Services (Hong Kong)
Senior DevOps Engineer- Financial Services (Hong Kong)

Randstad Hong Kong Limited • Hong Kong

On-site
HKD 900,000 - 1,200,000
System Analyst (Kubernetes, SRE) 65K
System Analyst (Kubernetes, SRE) 65K

Michael Page International (HK) Ltd • Hong Kong

On-site
HKD 720,000 - 1,200,000
Senior SRE & Platform Innovation Lead
Senior SRE & Platform Innovation Lead

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000
Core Site Reliability Engineer
Core Site Reliability Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,200,000
Senior System Engineer- DevOps (Redhat OpenShift | Finance)
Senior System Engineer- DevOps (Redhat OpenShift | Finance)

Manpower Services (Hong Kong) Limited • Hong Kong

On-site
HKD 700,000 - 1,100,000