Application Reliability Engineer

IO TECH SOLUTIONS LIMITED

Hong Kong

On-site

HKD 1,200,000 - 2,000,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IO TECH SOLUTIONS LIMITED is seeking a world-class Reliability Engineer to join a premier High-Frequency Trading firm. You will sit with traders and core engineers, engineering resiliency into ultra-low-latency systems and globally distributed clusters.

The role emphasizes proactive incident management, threat detection, and rapid recovery in a 24/7 environment. You will own the incident lifecycle, coordinate responders, and push teams to meet rigorous SLAs while maintaining observability and an

Qualifications

  • Production Operations, SRE, NOC/Command Centre, or Trading Operations experience, ideally in HFT or latency-sensitive environments.
  • Ability to coordinate major incidents and guide senior stakeholders to resolution.
  • Strong monitoring and observability skills; define service levels and SLAs.
  • Linux and networking fundamentals; read dashboards and logs to identify anomalies.
  • Exposure to containerized, cloud-hosted, and bare-metal systems (Kubernetes, Docker, GCP) is highly desirable.

Responsibilities

  • Proactive resilience: automate triage, alert enrichment, routing, and correlation to speed up first-line response.
  • Own incident lifecycle: assess impact, assemble responders, and manage resolution under a global framework.
  • Read symptoms quickly and stabilize systems using runbooks; escalate when needed.
  • Coordinate handover across EMEA, AMER, and APAC within a unified incident standard.

Skills

Incident management
SRE
Linux
Networking
Monitoring

Tools

PagerDuty
Jira Service Management
Grafana
Prometheus
Log aggregation

Job description

The Opportunity

We are seeking a world-class Reliability Engineer to join a premier High-Frequency Trading (HFT) firm. In our world, we measure success in microseconds and nanoseconds. Downtime isn't just a ticket—it's a direct, measurable hit to P&L by the minute.

This is not a conventional "keeping the lights on" role. You will sit shoulder-to-shoulder with traders, quantitative researchers, and core systems engineers, acting as the critical linchpin that keeps the global trading engine firing on all cylinders. You won't just react to problems; you will actively engineer resiliency into the fabric of one of the fastest trading environments on the planet.

Why You'll Love This Role
  • Massive P&L Impact: Your decisions directly protect (and unlock) millions in daily revenue. Every second of uptime you preserve is a tangible win for the firm.
  • Elite Compensation: We pay at the top of the market to attract the best. Your base salary and performance-based bonuses reflect the critical nature of this role.
  • Unmatched Autonomy: You own the room. As Incident Commander, your decisions hold authority—even when the call is filled with senior engineers, quants, or managing directors. You coordinate, delegate, and dictate the strategy.
  • Cutting-Edge Complexity: Manage ultra-low-latency architectures, globally distributed Kubernetes clusters, and highly advanced observability stacks at a scale and speed that few firms can match.
  • Zero Bureaucracy: We operate a flat structure. You have the standing to push back on development teams, infrastructure leads, or traders when operational standards slip.
What You Will Do

Proactive Resilience:

  • Automate the repetitive parts of triage (alert enrichment, routing, and correlation) so your first-line responders are 10x faster.
  • Obsess over monitoring gaps. If it can't be observed, it can't be traded. You will define service levels and push teams to meet rigorous SLAs.

Command the Response:

  • Take full control when things break. You assess the impact, assemble the right responders, and run the entire incident lifecycle under our Global Incident Management framework.
  • Use your deep technical breadth (Linux, Networking, Logs) to read symptoms instantly, stabilize systems using runbooks, and escalat cleanly when issues exceed documented steps.

Global Ownership:

  • Seamlessly hand over between EMEA, AMER, and APAC under one unified incident standard. You are part of a 24/7 elite global force.
What You Need to Succeed
  • Proven Experience: Background in Production Operations, SRE, NOC/Command Centre, or Trading Operations—ideally within HFT, financial services, or other extreme latency-sensitive environments.
  • Command Presence: A track record of coordinating major incidents. You aren't afraid to take the microphone and guide a room of senior stakeholders toward resolution.
  • Elite Triage Skills: You cut through assumptions under pressure. You know when to push forward and exactly when to pull in a specialist.
  • Technical Breadth (Not Just Depth): You are dangerous enough across all domains (Apps, Infrastructure, Data, Connectivity) to be useful everywhere.
  • Solid Fundamentals: Strong Linux and networking knowledge. You can read a dashboard, parse a log file, and spot the anomaly in seconds.
  • Tooling Mastery: Hands‑on with PagerDuty, Jira Service Management, Grafana, Prometheus, and log aggregation tools.
  • Automation Mindset: Scripting proficiency (Python preferred; Bash/Go are a bonus) applied to operational workflows—not just product code.
  • Bonus: Exposure to containerized, cloud-hosted, and bare-metal production systems (Kubernetes, Docker, GCP) is highly desirable.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Core Site Reliability Engineer
Core Site Reliability Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,200,000
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

Nahc • Hong Kong

On-site
HKD 900,000 - 1,200,000
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

Not Another Headhunting Company • Hong Kong

On-site
Senior DevOps / SRE Engineer
Senior DevOps / SRE Engineer

Red Begonia Limited • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
Application Support Engineer - Digital Assets - Hong Kong
Application Support Engineer - Digital Assets - Hong Kong

BAH Partners • Hong Kong

On-site
HKD 420,000 - 680,000
Competitive compensation
Comprehensive benefits
Global exposure
Core DevOps Engineer
Core DevOps Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,300,000
Relocation support
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

nahc.io • Hong Kong

On-site
HKD 391,000 - 627,000
Senior DevOps Engineer
Senior DevOps Engineer

Aurosglobal • Hong Kong

On-site
HKD 800,000 - 1,200,000
Flexible working environment
Competitive compensation
Ownership of core infrastructure
Sr. Manager, Site Reliability & Innovation, IT
Sr. Manager, Site Reliability & Innovation, IT

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000
Infrastructure Operations Engineer - Banking
Infrastructure Operations Engineer - Banking

IO TECH SOLUTIONS LIMITED • Hong Kong

On-site
HKD 400,000 - 640,000