Site Reliability Engineer, Data & Observability Platform

Pulsar

Hong Kong

On-site

HKD 400,000 - 700,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Pulsar in Hong Kong is seeking an experienced Site Reliability Engineer focused on observability and streaming data pipelines. You will own the platform’s end-to-end monitoring, instrument services, build Grafana dashboards and alerts that clearly indicate what broke and what to check next, and participate in on-call rotations.

The role requires deploying and running production systems on Kubernetes (EKS), managing infrastructure as code with Terraform, and delivering changes through CI/CD and

Qualifications

  • University degree in CS, Software Engineering, or related field.
  • 4-7 years of relevant working experience in SRE/Observability or data platforms.
  • Experience instrumenting services, building dashboards, writing actionable alerts, and diagnosing live incidents.
  • Experience with CI/CD pipelines and GitOps practices.
  • Experience with distributed stream processing engines and OLAP databases.
  • Strong programming in at least one of Python, Scala, Java or Rust, plus SQL skills.

Responsibilities

  • Own the observability of the platform end to end: decide what to measure, instrument it, collect it, store it, and present it to people.
  • Build Grafana dashboards and alerts that name what broke and what to check next.
  • Run production systems: diagnose incidents, recover data, and fix root causes to prevent recurrence.
  • Ship changes through CI/CD and GitOps; understand the difference between merged and live.
  • Operate streaming workloads on Kubernetes (EKS) with capacity planning and config delivery.
  • Provision AWS resources behind the platform as code (Terraform).
  • Participate in on-call for the data and observability platform.
  • Build and maintain near-realtime pipelines carrying telemetry and logs to quarriable tables.
  • Reason about delivery guarantees and their impact on dashboard numbers.
  • Design data quality checks to route bad records for investigation.
  • Design and optimize OLAP tables for fast billions-row queries.
  • Write and tune SQL behind dashboards and analytics.
  • Translate requests into pipeline changes and table designs.

Skills

Observability
Grafana
Kubernetes
Terraform
SQL
Programming (Python/Scala/Java/Rust)
CI/CD
Streaming
OLAP/Columnar
On-call

Education

Bachelor's degree in Computer Science or related

Tools

Prometheus
ELK/OpenTelemetry
CloudWatch

Job description

  • You are pioneering and innovative and want to be part of the cutting-edge and disruptive crypto-currency world
  • You are eager to learn new knowledge in both financial and technical fields
  • You thrive in a non-hierarchical organization with a casual working environment
  • You enjoy solving complex distributed systems challenges and optimizing streaming data pipelines
  • You value comprehensive documentation and collaborative problem-solving
Team / Role

As an Engineer you will:

SRE & Observability
  • Own the observability of the platform end to end: decide what to measure, instrument it, collect it, store it, and put it in front of people
  • Build Grafana dashboards and alerts that name what broke and what to check next — not alerts that merely report a number crossing a line
  • Run the production systems: diagnose silent failures, recover missed data, and fix the cause so it doesn’t recur
  • Ship changes through CI/CD and GitOps, and know the difference between “merged” and “live”
  • Operate streaming workloads on Kubernetes (EKS) — resource budgeting, config delivery, capacity, and why an OOM-kill can also be a data-loss event
  • Provision the AWS resources behind the platform as code (we use Terraform; we’ll show you ours)
  • Participate in on-call for the data and observability platform
  • Build and maintain the near-realtime pipelines that carry telemetry and logs from source to quarriable table
  • Reason about delivery guarantees (at-most-once vs at-least-once) and what they mean for the numbers on a dashboard
  • Design data quality checks — route bad records for investigation rather than dropping them silently
  • Design and optimize OLAP tables so billion-row queries return fast: sort keys, partitioning, deduplication, retention
  • Write and tune the SQL behind dashboards and analytics
  • Translate a request like “show me p99 latency per venue” into a pipeline change and a table design
Required Skillset
  • University degree in Computer Science, Software Engineering or related disciplines
  • At least 4-7 years relevant working experience
  • Observability and monitoring in production — instrumenting services, building dashboards, writing alerts that are worth acting on, and diagnosing a live incident from metrics and logs together (Prometheus, Grafana, Loki, ELK, Datadog, CloudWatch — any equivalent stack)
  • CI/CD — delivering your own changes through a pipeline (GitLab CI, Jenkins, GitHub Actions, or similar), not handing them to someone else
  • A distributed stream processing engine in production (Flink, Spark Structured Streaming, or similar)
  • An OLAP / columnar database — designing tables and tuning queries, not just reading from them (ClickHouse, BigQuery, Druid, Redshift, or similar)
  • Kubernetes working knowledge — deployments, resource requests and limits, config delivery, reading pod state
  • Strong programming in at least one of Python, Scala, Java or Rust, plus fluent analytical SQL
  • Ability to troubleshoot production systems methodically, and the instinct to fix the cause rather than the symptom
It is a plus if:
  • OpenTelemetry, distributed tracing, or high-cardinality telemetry at scale
  • Airflow (DAGs, scheduling, task orchestration)
  • Infrastructure as code beyond the basics (Terraform modules, drift, targeted applies)
  • Build tools for JVM or multi-module projects (Maven, Gradle, sbt, Cargo)
  • Familiarity with message queues (SQS, Kafka, or similar) and event-driven architec

We encourage applicants to read through our Privacy Notice for Applicants before submitting your applications - https://pulsar.com/careers/

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - HFT
Site Reliability Engineer - HFT

Selby Jennings • Hong Kong

On-site
HKD 480,000 - 720,000
Site Reliability Engineer - J13203
Site Reliability Engineer - J13203

Pinpoint Asia • Hong Kong

On-site
HKD 600,000 - 900,000
SRE for Real-Time Streaming & Observability
SRE for Real-Time Streaming & Observability

Pulsar • Hong Kong

On-site
HKD 400,000 - 700,000
Mid-Level SRE
Mid-Level SRE

IO TECH SOLUTIONS LIMITED • Hong Kong

On-site
HKD 480,000 - 600,000
Sr. Manager, Site Reliability & Innovation, IT
Sr. Manager, Site Reliability & Innovation, IT

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000
Backend/Systems Engineer (Real-Time System)
Backend/Systems Engineer (Real-Time System)

Nahc • Hong Kong

On-site
HKD 900,000 - 1,200,000
Site Reliability Engineer – SRE / Data & Platform
Site Reliability Engineer – SRE / Data & Platform

IO TECH SOLUTIONS LIMITED • Hong Kong

On-site
HKD 600,000 - 1,000,000
Senior Site Reliability Engineer (APAC)
Senior Site Reliability Engineer (APAC)

Reap • Hong Kong

On-site
HKD 900,000 - 1,500,000
Application Reliability Engineer
Application Reliability Engineer

IO Tech Solutions Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
Site Reliability Engineer: Observability & Data Pipelines
Site Reliability Engineer: Observability & Data Pipelines

Pinpoint Asia • Hong Kong

On-site
HKD 600,000 - 900,000