Site Reliability Engineer – SRE / Data & Platform

IO TECH SOLUTIONS LIMITED

Hong Kong

On-site

HKD 600,000 - 1,000,000

Full time

17 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

IO TECH SOLUTIONS LIMITED is seeking a Site Reliability Engineer to strengthen production reliability, observability and performance of critical systems.

The role requires hands-on experience with data pipelines, streaming or data platforms, and proficiency in Python or scripting, plus Kubernetes and cloud environments.

Qualifications

  • 3–7 years of experience in SRE/DevOps or related role.
  • Hands-on experience with production systems, monitoring, and incident troubleshooting.
  • Must have practical exposure to data-related infrastructure: data pipelines, streaming, ETL/ELT, Kafka/Kinesis, Spark/Flink, data platforms or data lakes.
  • Experience with Kubernetes and cloud environments (AWS, GCP or Azure).
  • Proficiency in Python or other scripting languages.
  • Strong communication and problem-solving skills.
  • Hong Kong-based candidates are highly preferred.

Responsibilities

  • Own and improve the reliability, availability and performance of production systems.
  • Build and enhance observability, monitoring, dashboards and alerting for critical services.
  • Troubleshoot production issues, perform root-cause analysis and drive reliability improvements.
  • Support Kubernetes/cloud-based infrastructure and production workloads.
  • Develop and maintain CI/CD and GitOps workflows.
  • Work with data pipelines, streaming workloads or data platform infrastructure to support reliable and scalable data processing.
  • Monitor data flow, system performance and potential issues such as latency, failures, capacity constraints and data loss.
  • Develop automation and operational tools using Python or other programming/scripting languages.
  • Participate in production support and on-call activities.

Skills

SRE/DevOps
Observability
Kubernetes
Python
CI/CD
Data pipelines
Cloud platforms
Incident troubleshooting

Tools

Kafka
Kinesis
Spark
Flink
OLAP databases

Job description

We are looking for a Site Reliability Engineer with strong production reliability and observability experience, together with hands-on exposure to data pipelines, streaming or data platforms.

You will work in a fast-paced technology environment supporting high-performance production systems, with a focus on reliability, monitoring, automation and system performance.

Key Responsibilities
  • Own and improve the reliability, availability and performance of production systems.
  • Build and enhance observability, monitoring, dashboards and alerting for critical services.
  • Troubleshoot production issues, perform root-cause analysis and drive reliability improvements.
  • Support Kubernetes/cloud-based infrastructure and production workloads.
  • Develop and maintain CI/CD and GitOps workflows.
  • Work with data pipelines, streaming workloads or data platform infrastructure to support reliable and scalable data processing.
  • Monitor data flow, system performance and potential issues such as latency, failures, capacity constraints and data loss.
  • Develop automation and operational tools using Python or other programming/scripting languages.
  • Participate in production support and on-call activities.
Requirements
  • 3–7 years of experience in SRE, DevOps, Platform Engineering, Production Engineering or a related role.
  • Strong hands-on experience with production systems, monitoring/observability and incident troubleshooting.
  • Must have practical exposure to data-related infrastructure, such as:
    • Data pipelines / data ingestion
    • Streaming platforms
    • ETL / ELT
    • Kafka / Kinesis or similar technologies
    • Spark / Flink or similar processing frameworks
    • Data platforms / data lakes / OLAP databases
  • Experience with Kubernetes and cloud environments such as AWS, GCP or Azure.
  • Experience with CI/CD, GitOps or infrastructure automation.
  • Proficiency in Python or another programming/scripting language.
  • Good understanding of system performance, reliability and troubleshooting.
  • Strong communication and problem-solving skills.
  • Hong Kong-based candidates are highly preferred.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Global Financial Institution - Hong Kong
Site Reliability Engineer - Global Financial Institution - Hong Kong

NLS Executive Search • Hong Kong

On-site
HKD 600,000 - 900,000
Very competitive compensation (base +
Site Reliability Engineer - J13203
Site Reliability Engineer - J13203

Pinpoint Asia • Hong Kong

On-site
HKD 600,000 - 900,000
Site Reliability Engineer - HFT
Site Reliability Engineer - HFT

Selby Jennings • Hong Kong

On-site
HKD 480,000 - 720,000
Sr. Manager, Site Reliability & Innovation, IT
Sr. Manager, Site Reliability & Innovation, IT

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000
Site Reliability Engineer — Observability & Data Platform
Site Reliability Engineer — Observability & Data Platform

Selby Jennings • Hong Kong

On-site
HKD 480,000 - 720,000
Site Reliability and Observability Engineer
Site Reliability and Observability Engineer

EXIO (HK) LIMITED • Hong Kong Island

On-site
HKD 450,000 - 700,000
SRE for Real-Time Streaming & Observability
SRE for Real-Time Streaming & Observability

Pulsar • Hong Kong

On-site
HKD 400,000 - 700,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

GFT TECHNOLOGIES SE • Hong Kong

On-site
HKD 700,000 - 1,200,000
Mid-Level SRE
Mid-Level SRE

IO TECH SOLUTIONS LIMITED • Hong Kong

On-site
HKD 480,000 - 600,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kody • Hong Kong

On-site
HKD 900,000 - 1,500,000
Competitive Package
A dynamic and innovative team
Collaborative, inclusive working env.