Reliability Engineer (Trading Platforms)

IO Tech Solutions Limited

Hong Kong

On-site

HKD 420,000 - 640,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IO Tech Solutions Limited seeks a Reliability Engineer for Trading Platforms in Hong Kong. You will automate triage workflows, improve alert quality, and track reliability across critical trading apps and dependencies.

Collaboration with traders, developers, and IT is essential to pinpoint service degradations. You will declare incidents as needed, act as Incident Commander, coordinate responders, and maintain clear incident timelines.

Qualifications

  • Production ops experience in latency-sensitive environment.
  • Strong triage and prioritization under pressure.
  • Clear communication to traders and engineers.

Responsibilities

  • Automate repeatable triage workflows (alert enrichment, routing, correlation).
  • Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
  • Track reliability and availability across critical trading applications and their dependencies; partner with users and IT.
  • Triage incoming alerts, issues, and escalations; assess impact, urgency, ownership.
  • Declare incidents when criteria are met; act as Incident Commander.
  • Coordinate responders and stakeholders; keep incident calls focused on facts and mitigation.
  • Maintain timelines, actions, and status updates throughout incident lifecycle.
  • Recover and stabilize systems using runbooks; escalate beyond documented steps.
  • Support post-incident reviews (PIR) and recurring issue reviews.
  • Ensure smooth handovers across EMEA, AMER, APAC with a single global model.

Skills

Production operations experience
SRE/Incidence handling
Triage and prioritization
Clear communication
Cross-domain collaboration
Linux fundamentals
Networking fundamentals
Python scripting
Kubernetes
Docker
Grafana
Prometheus
PagerDuty
Jira Service Management
Log analysis
GCP/cloud familiarity

Tools

Kubernetes
Docker
Grafana
Prometheus
PagerDuty
Jira Service Management
Linux
Networking tools
GCP

Job description

Reliability Engineer (Trading Platforms)
  • Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g., alert enrichment, routing, correlation, and operational runbooks).
  • Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
  • Track reliability and availability across critical trading applications and their dependencies. Partner with users, development teams, and IT to pinpoint where service levels are degrading.
  • Triage incoming alerts, issues, and escalations—assessing impact, urgency, and ownership.
  • Determine when incident criteria are met, declare incidents, and act as Incident Commander.
  • Coordinate responders and stakeholders; keep incident calls focused on facts, mitigation, and recovery.
  • Maintain clear timelines, actions, and status updates throughout the incident lifecycle.
  • Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
  • Support post-incident review (PIR) follow-ups and recurring issue reviews.
  • Ensure smooth handovers across EMEA, AMER, and APAC using a single global model: one incident standard and one handover process.

Requirements:

  • Experience in production operations, SRE, NOC/command center, trading operations, or a comparable first-line technical role—ideally in a trading, financial services, or other latency-sensitive environment.
  • Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
  • Clear communication (verbal and written): status updates are understandable to both traders and engineers.
  • Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
  • Solid Linux and networking fundamentals, plus the ability to quickly interpret alerts, logs, dashboards, and symptoms.
  • Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data).
  • Familiarity with incident and observability tooling (e.g., PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search).
  • Scripting/automation skills (Python preferred; Bash and Go are a plus), applied to triage, enrichment, routing, and correlation (not product code).
  • Exposure to containerized/cloud-hosted production environments (Kubernetes, Docker, GCP) is a plus.

EnvironmentHandoverGoAPACDataRoutingConnectivityPrometheusMitigationStepsSupportSearchDevelopmentInterfacesGrafanaFinancial ServicesGcpOperationsTradingOwnershipBashTimelinesReviewsReliabilityAvailabilityInfrastructureNetworkingAutomationKubernetesPressureLinuxDockerJIRAPythonCommunicationManagement

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Application Reliability Engineer
Application Reliability Engineer

IO Tech Solutions Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
Core Site Reliability Engineer
Core Site Reliability Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,200,000
Site Reliability Engineer, Observability, Investment Bank
Site Reliability Engineer, Observability, Investment Bank

Recruit Logic Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
Trading Platform Reliability Engineer
Trading Platform Reliability Engineer

IO Tech Solutions Limited • Hong Kong

On-site
HKD 420,000 - 640,000
Application Support Engineer (Strong Trading System Experience)
Application Support Engineer (Strong Trading System Experience)

Luxoft • Hong Kong

On-site
HKD 480,000 - 720,000
Infrastructure & Market Connectivity Engineer (Trading and Brokerage Industry)
Infrastructure & Market Connectivity Engineer (Trading and Brokerage Industry)

Alexis Services Limited • Hong Kong

On-site
HKD 350,000 - 650,000
Production Reliability Engineer / SRE (Hong Kong)
Production Reliability Engineer / SRE (Hong Kong)

IO Tech Solutions Limited • Hong Kong

On-site
HKD 720,000 - 960,000
Technology Support Lead, Prime Finance and Clearing Production Management Support
Technology Support Lead, Prime Finance and Clearing Production Management Support

JPMorgan Chase & Co. • Hong Kong

On-site
HKD 800,000 - 1,200,000
SL2 Technical Support Engineer
SL2 Technical Support Engineer

Luxoft • Hong Kong

On-site
HKD 480,000 - 840,000
Application Production Support Engineer
Application Production Support Engineer

Newtone consulting • Hong Kong

On-site
HKD 450,000 - 700,000