Reliability Engineer

Flow Traders

Hong Kong

On-site

HKD 600,000 - 900,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Flow Traders is seeking Reliability Engineers to safeguard the production performance of our global trading platform and the technology estate around it. You will automate triage, lead incident responses, and coordinate across traders, developers, and IT as part of our Global Incident Management framework.

Ideal candidates have hands-on experience in production operations or SRE, strong guidance in incident handling, and a proven ability to push for improved monitoring and runbooks.

Qualifications

  • Experience in production operations, SRE, NOC/command centre or similar first-line role.
  • Track record coordinating major incidents and leading calls with senior stakeholders.
  • Strong triage, prioritisation and clear communication under time pressure.

Responsibilities

  • Automate repeatable triage work to speed up first-line responses.
  • Push for monitoring, alerting and visibility to close gaps.
  • Track reliability and availability across trading apps and platforms.
  • Call out weak ownership and drive fixes with owning teams.
  • Triage alerts, assess impact, urgency and ownership.
  • Declare incidents when criteria are met and act as Incident Commander.
  • Coordinate responders, keep calls focused on facts and recovery.
  • Maintain timelines, actions and status updates during incidents.
  • Recover or stabilise systems using runbooks; handover per process.
  • Perform cross-team operational tasks as needed.
  • Support PIR follow-up and recurring issue reviews.
  • Hand over cleanly between EMEA, AMER and APAC under global model.

Skills

Production operations
SRE
NOC/command centre
Trading operations
Incident management
Communication
Judgment and escalation

Tools

PagerDuty
Jira Service Management
Grafana
Prometheus
Log search
Python
Kubernetes
Docker
GCP

Job description

Flow Traders is hiring Reliability Engineers to safeguard the production performance of our global trading platform and the technology estate around it. Undetected degradation costs P&L by the minute.

Most of what makes failures smaller, shorter and easier to contain happens before anything breaks. You automate the repetitive parts of triage, keep alerts to the ones that need a human, and close monitoring gaps. Root causes go back to the teams that own them.

When something does break, control of the room is yours. You're first on it: you size it up, decide who is on the call, and run the response as Incident Commander under our Global Incident Management framework. You coordinate and delegate rather than dropping into the debugging, and your decisions hold even when the room is more senior than you.

You'll sit alongside traders, developers, infrastructure and specialist teams, with the standing to push any of them when operational standards slip.

What You Will Do
Making failures smaller
  • Automate repeatable triage work so first-line responds faster and more consistently, including alert enrichment, routing, correlation and operational workflows
  • Push for monitoring, alerting and visibility where gaps exist
  • Track reliability and availability across critical trading applications and the platforms they depend on, and work with users, development teams and IT to find where service levels are degrading
  • Call out weak ownership, missed SLAs, poor alerts and ineffective runbooks, and drive the owning teams to fix them
Running the response
  • Triage incoming alerts, issues and escalations, and assess impact, urgency and ownership
  • Decide when incident criteria are met, declare the incident, and act as Incident Commander
  • Coordinate responders and stakeholders, and keep incident calls focused on facts, mitigation and recovery
  • Maintain clear timelines, actions and status updates throughout an incident
  • Recover or stabilise systems using approved runbooks, and escape cleanly through the defined support and development path when the issue goes beyond documented recovery steps
  • Perform common operational tasks across adjacent teams where needed
After the incident
  • Support PIR follow-up and recurring issue review
  • Hand over cleanly between EMEA, AMER and APAC under one global model, one incident standard, one handover process
What You Will Need To Succeed
The role
  • Experience in production operations, SRE, NOC/command centre, trading operations or a similar first-line technical role, ideally in a trading, financial services or other latency-sensitive environment
  • Track record of running or coordinating major incidents, and comfort taking command of a call with senior people on it
  • Strong triage and prioritisation. You can separate facts from assumptions under time pressure and keep the response moving
  • Clear verbal and written communication. Your status updates are readable by a trader and an engineer at the same time
  • Strong judgment and escalation discipline. You know when to keep going and when to pull in a specialist
  • Willingness to hold the line on process, and to push back when poor operational behaviour creates risk for trading
The technology
  • Technically broad rather than deep. You need enough understanding of how most teams operate to be useful across domains, not to be the specialist resolver
  • Solid Linux and networking fundamentals, and the ability to read alerts, logs, dashboards and symptoms quickly
  • Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data)
  • Familiarity with incident and observability tooling: PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search
  • Scripting and automation ability (Python preferred; Bash, Go a plus) applied to triage, enrichment, routing and correlation rather than to product code
  • Exposure to containerised and cloud-hosted production systems (Kubernetes, Docker, GCP) is a plus
What Good Looks Like

Six months in, a strong Reliability Engineer here has already made the next incident smaller. They've automated something that used to be manual, retired alerts nobody could act on, and closed a monitoring gap that was costing us detection time. When something does break they are calm under pressure, clear in communication and disciplined in process. They can control a noisy incident without trying to become the specialist resolver, and they know how to use a runbook safely and when to And they don't let poor operational standards slide when trading is exposed.

Flow Traders does not accept unsolicited resumes from any professional staffing or search firms. All resumes, and any other information identifying potential candidates, submitted to any employee at Flow Traders via-email, the Internet or directly without a valid and signed search agreement will be deemed free to contact by Flow Traders without any restrictions and no placement fee of any kind will be paid in the event the candidate is hired by Flow Traders.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Reliability Engineer & Incident Commander
Reliability Engineer & Incident Commander

Flow Traders • Hong Kong

On-site
HKD 600,000 - 900,000
Application Reliability Engineer
Application Reliability Engineer

IO Tech Solutions Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
Core Site Reliability Engineer
Core Site Reliability Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,200,000
Trading Systems Reliability Engineer - C++
Trading Systems Reliability Engineer - C++

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,300,000
Trading production engineer
Trading production engineer

CW Talent Solutions • Hong Kong

On-site
HKD 800,000 - 1,100,000
Site Reliability Engineer - HFT
Site Reliability Engineer - HFT

Selby Jennings • Hong Kong

On-site
HKD 480,000 - 720,000
Operations Specialist
Operations Specialist

Flow Traders • Hong Kong

On-site
HKD 400,000 - 600,000
Senior Research Engineer
Senior Research Engineer

Flow Traders • Hong Kong

On-site
HKD 800,000 - 1,000,000
Senior Software Engineer, C++
Senior Software Engineer, C++

Flow Traders • Hong Kong

On-site
HKD 700,000 - 900,000
Senior Software Engineer, C++
Senior Software Engineer, C++

Flow Traders • Hong Kong

On-site
HKD 900,000 - 1,300,000