Application Reliability Engineer

IO Tech Solutions Limited

Hong Kong

On-site

HKD 900,000 - 1,500,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IO Tech Solutions Limited in Hong Kong is seeking a world-class Reliability Engineer to join a premier High-Frequency Trading firm. You will work shoulder-to-shoulder with traders, quantitative researchers, and core systems engineers to keep the global trading engine resilient.

This role requires proactive incident management, deep technical breadth across Linux, networking, and observability, and the autonomy to drive uptime in a 24/7 environment.

Qualifications

  • Proven experience coordinating major incidents in high-pressure environments.
  • Strong Linux and networking fundamentals with the ability to read dashboards and logs quickly.
  • Experience with observability stacks and incident management processes.

Responsibilities

  • Automate triage processes to speed up first-line response.
  • Take ownership of incident lifecycle from detection to resolution.
  • Coordinate with global teams across multiple time zones to restore service.
  • Maintain and improve monitoring, alerting, and runbooks for critical systems.

Skills

Production Operations
SRE
NOC/Command Centre
Trading Operations
Linux
Networking

Tools

PagerDuty
Jira Service Management
Grafana
Prometheus
Log aggregation tools
Kubernetes
Docker

Job description

We are seeking a world-class Reliability Engineer to join a premier High-Frequency Trading (HFT) firm. In our world, we measure success in microseconds and nanoseconds. Downtime isn't just a ticket—it's a direct, measurable hit to P&L by the minute.

This is not a conventional "keeping the lights on" role. You will sit shoulder-to-shoulder with traders, quantitative researchers, and core systems engineers, acting as the critical linchpin that keeps the global trading engine firing on all cylinders. You won't just react to problems; you will actively engineer resiliency into the fabric of one of the fastest trading environments on the planet.

The Opportunity

We are seeking a world-class Reliability Engineer to join a premier High-Frequency Trading (HFT) firm. In our world, we measure success in microseconds and nanoseconds. Downtime isn't just a ticket—it's a direct, measurable hit to P&L by the minute.

This is not a conventional "keeping the lights on" role. You will sit shoulder-to-shoulder with traders, quantitative researchers, and core systems engineers, acting as the critical linchpin that keeps the global trading engine firing on all cylinders. You won't just react to problems; you will actively engineer resiliency into the fabric of one of the fastest trading environments on the planet.

Why You’ll Love This Role
  • Massive P&L Impact: Your decisions directly protect (and unlock) millions in daily revenue. Every second of uptime you preserve is a tangible win for the firm.
  • Elite Compensation: We pay at the top of the market to attract the best. Your base salary and performance-based bonuses reflect the critical nature of this role.
  • Unmatched Autonomy: You own the room. As Incident Commander, your decisions hold authority—even when the call is filled with senior engineers, quants, or managing directors. You coordinate, delegate, and dictate the strategy.
  • Cutting-Edge Complexity: Manage ultra-low-latency architectures, globally distributed Kubernetes clusters, and highly advanced observability stacks at a scale and speed that few firms can match.
  • Zero Bureaucracy: We operate a flat structure. You have the standing to push back on development teams, infrastructure leads, or traders when operational standards slip.
What You Will Do
Proactive Resilience:
  • Automate the repetitive parts of triage (alert enrichment, routing, and correlation) so your first-line responders are 10x faster.
  • Obsess over monitoring gaps. If it can't be observed, it can't be traded. You will define service levels and push teams to meet rigorous SLAs.
Command the Response:
  • Take full control when things break. You assess the impact, assemble the right responders, and run the entire incident lifecycle under our Global Incident Management framework.
  • Use your deep technical breadth (Linux, Networking, Logs) to read symptoms instantly, stabilize systems using runbooks, and elevate cleanly when issues exceed documented steps.
Global Ownership:
  • Seamlessly hand over between EMEA, AMER, and APAC under one unified incident standard. You are part of a 24/7 elite global force.
What You Need to Succeed
  • Proven Experience: Background in Production Operations, SRE, NOC/Command Centre, or Trading Operations—ideally within HFT, financial services, or other extreme latency-sensitive environments.
  • Command Presence: A track record of coordinating major incidents. You aren't afraid to take the microphone and guide a room of senior stakeholders toward resolution.
  • Elite Triage Skills: You cut through assumptions under pressure. You know when to push forward and exactly when to pull in a specialist.
  • Technical Breadth (Not Just Depth): You are dangerous enough across all domains (Apps, Infrastructure, Data, Connectivity) to be useful everywhere.
  • Solid Fundamentals: Strong Linux and networking knowledge. You can read a dashboard, parse a log file, and spot the anomaly in seconds.
  • Tooling Mastery: Hands-on with PagerDuty, Jira Service Management, Grafana, Prometheus, and log aggregation tools.
  • Automation Mindset: Scripting proficiency (Python preferred; Bash/Go are a bonus) applied to operational workflows—not just product code.
  • Bonus: Exposure to containerized, cloud-hosted, and bare-metal production systems (Kubernetes, Docker, GCP) is highly desirable.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer (Trading Platforms)
Reliability Engineer (Trading Platforms)

IO Tech Solutions Limited • Hong Kong

On-site
HKD 420,000 - 640,000
Core Site Reliability Engineer
Core Site Reliability Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,200,000
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

Nahc • Hong Kong

On-site
HKD 900,000 - 1,200,000
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

Not Another Headhunting Company • Hong Kong

On-site
Application Support Engineer - Digital Assets - Hong Kong
Application Support Engineer - Digital Assets - Hong Kong

BAH Partners • Hong Kong

On-site
HKD 420,000 - 680,000
Competitive compensation
Comprehensive benefits
Global exposure
Core DevOps Engineer
Core DevOps Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,300,000
Relocation support
Senior DevOps Engineer
Senior DevOps Engineer

Aurosglobal • Hong Kong

On-site
HKD 800,000 - 1,200,000
Flexible working environment
Competitive compensation
Ownership of core infrastructure
Infrastructure Operations Engineer - Banking
Infrastructure Operations Engineer - Banking

IO Tech Solutions Limited • Hong Kong

On-site
HKD 600,000 - 800,000
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

nahc.io • Hong Kong

On-site
HKD 391,000 - 627,000
Sr. Manager, Site Reliability & Innovation, IT
Sr. Manager, Site Reliability & Innovation, IT

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000