Sr. Manager, Site Reliability & Innovation, IT

CLSA

Hong Kong

On-site

HKD 900,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CLSA in Hong Kong is seeking a Senior Site Reliability Engineer to own monitoring, SRE operations, and the stability of critical systems. You will design and run end-to-end observability across distributed environments, working to improve reliability and performance.

The role requires strong Linux administration, hands-on experience with Prometheus, Grafana, Elasticsearch, Kibana, and Kubernetes, plus automation with Bash, Python, Ansible, and CI/CD pipelines.

Qualifications

  • Bachelor’s degree or higher in Computer Science / Engineering.

Responsibilities

  • Own monitoring, Kubernetes platform reliability, and SRE operations to ensure highly reliable, available, and performant systems
  • Build, enhance, and maintain monitoring solutions using ITRS Geneos, Prometheus, Victoria-Metrics, Elasticsearch, and Grafana
  • Develop, optimize, and maintain alerting rules, dashboards, and observability pipelines
  • Troubleshoot Linux servers (RHEL 7/8/9), including upgrades, configurations, patching, and maintenance, while determining appropriate monitoring requirements for system changes
  • Analyze logs, investigate issues, and perform fault finding to identify performance exceptions
  • Collaborate with engineering, application, and infrastructure teams to improve system resilience, stability, security, efficiency, and scalability
  • Operate, maintain, and optimize Kubernetes environments, including cluster health, workload reliability, capacity planning, and platform observability
  • Continuously research and adopt modern monitoring and SRE tools and practices.

Skills

Monitoring platforms
Kubernetes
Linux administration
Prometheus
Grafana
Elasticsearch
Kibana
ITRS Geneos
VictoriaMetrics
Automation
CI/CD
Python
Bash

Education

Bachelor’s degree or higher in Computer Science / Engineering

Tools

ITRS Geneos
Prometheus
VictoriaMetrics
Elasticsearch
Grafana
Kibana
Logstash
Kubernetes
Bash
Python
Ansible
CI/CD tools
Exporters

Job description

We are seeking a Senior Site Reliability Engineer who will be responsible for both build and shared services operations, including monitoring, site reliability engineering (SRE), and ensuring the stability, scalability, and performance of critical systems.

The ideal candidate is a strong technical problem-solver and capable of delivering end-to-end monitoring and reliability solutions while diagnosing complex issues during critical incidents.

Key Areas of Responsibilities

  • Own monitoring, Kubernetes platform reliability, and SRE operations to ensure highly reliable, available, and performant systems
  • Build, enhance, and maintain monitoring solutions using ITRS Geneos, Prometheus, Victoria-Metrics, Elasticsearch, and Grafana
  • Develop, optimize, and maintain alerting rules, dashboards, and observability pipelines
  • Troubleshoot Linux servers (RHEL 7/8/9), including upgrades, configurations, patching, and maintenance, while determining appropriate monitoring requirements for system changes
  • Analyze logs, investigate issues, and perform fault finding to identify performance exceptions
  • Collaborate with engineering, application, and infrastructure teams to improve system resilience, stability, security, efficiency, and scalability.
  • Operate, maintain, and optimize Kubernetes environments, including cluster health, workload reliability, capacity planning, and platform observability
  • Continuously research and adopt modern monitoring and SRE tools and practices.

Requirements

  • Bachelor’s degree or higher in Computer Science / Engineering
  • Around 8-10 years of experience within IT, preferably in site reliability engineering, production support, platform engineering, or investment banking environments
  • Strong experience configuring and maintaining monitoring and observability platforms, including:
  • ITRS Geneos, Prometheus, Victoriametrics, Elasticsearch, Grafana, and Kibana
  • Experience with automation (e.g., Bash, Python, Ansible, CI/CD tools) is a must
  • Hands-on experience building and implementing Prometheus pipelines, including exporters, scraping configurations, relabelling, metric routing, and integrations with long-term storage (e.g., Victoriametrics)
  • Experience building and maintaining Logstash pipelines, including ingestion, parsing, filtering, enrichment, and routing of logs into Elasticsearch
  • Ability to design, build, and maintain Grafana and Kibana dashboards for metrics, logs, and performance analytics across distributed systems
  • Understanding of metrics, logging, alerting, dashboards, and observability pipelines
  • Strong Linux administration skills (RHEL 7/8/9), including troubleshooting, upgrades, configuration, patching, and performance optimization.
  • Good understanding of SRE principles, high availability, scalability, incident management and Disaster Recovery / Business Continuity Planning) activities
  • Experience managing GPU-enabled infrastructure for AI or machine learning platforms is preferred.
  • Strong hands-on experience with Kubernetes, including cluster operations, workload orchestration, troubleshooting, scaling, and production support
  • Understanding of networking fundamentals, performance tuning, and troubleshooting distributed systems
  • Operations with participation in on-call rotations, including after-hours and weekend support
  • Self-motivated, adaptable and able to prioritize, learn continuously and manage multiple responsibilities effectively
  • Excellent in English, with Chinese will be advantage
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

TEKsystems • Hong Kong

On-site
HKD 500,000 - 900,000
Senior SRE & Platform Innovation Lead
Senior SRE & Platform Innovation Lead

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000
DevOps / Site Reliability Engineer
DevOps / Site Reliability Engineer

Talentquest HR Limited • Hong Kong

On-site
HKD 420,000 - 640,000
System Analyst - Kubernetes 65K
System Analyst - Kubernetes 65K

Michael Page International (Hong Kong) Limited • Hong Kong

On-site
HKD 614,000 - 1,004,000
Competitive salary
Benefits package
System Analyst (Kubernetes, SRE) 65K
System Analyst (Kubernetes, SRE) 65K

Michael Page International (HK) Ltd • Hong Kong

On-site
HKD 720,000 - 1,200,000
System Analyst - Financial Services (Work Life Balance)
System Analyst - Financial Services (Work Life Balance)

TEKsystems • Hong Kong

On-site
HKD 420,000 - 660,000
Work Life Balance
Devops Engineer or Software Engineer or System Analyst
Devops Engineer or Software Engineer or System Analyst

Recruit Squad Limited • Hong Kong

On-site
HKD 420,000 - 640,000
SRE/ DevOps Engineer- Financial Services
SRE/ DevOps Engineer- Financial Services

Captar Partners Limited • Hong Kong

On-site
HKD 700,000 - 1,100,000
Senior SRE & DevOps Engineer: Platform Reliability
Senior SRE & DevOps Engineer: Platform Reliability

TEKsystems • Hong Kong

On-site
HKD 500,000 - 900,000
Application Support Engineer - Digital Assets - Hong Kong
Application Support Engineer - Digital Assets - Hong Kong

BAH Partners • Hong Kong

On-site
HKD 420,000 - 680,000
Competitive compensation
Comprehensive benefits
Global exposure