A complete application in a minute — tailored resume and cover letter, ready to send.
EXIO (HK) LIMITED is seeking a Site Reliability and Observability Engineer to build and maintain observability across on-prem and cloud systems.
You will drive reliability improvements through performance analysis, capacity planning, stress tests, and incident response; fine-tune alerts, validate monitoring after releases, and define reliability targets (SLA/RTO) for critical services.
Build and maintain observability using logs, metrics, and dashboards for on-prem and cloud systems.
Drive reliability improvements with performance analysis, capacity planning, stress tests, restore drills, failure prevention.
Fine-tune alerts to reduce noise and improve signal quality and response speed; validate monitoring/alerts after releases and during deployments.
Define and track reliability targets (SLA/RTO) for critical services.
Improve incident response by maintaining operational procedures, service catalogs and clear escalation paths.
Automate routine operational tasks e.g. health checks, remediation, validation gates.
Perform Root Cause Analysis for reliability incidents and implement preventative actions.
Ensure observability configuration changes are controlled and audit-evidenced.
Uphold IT General Control and compliance standard with evidence retention, access controls, change approvals.
Requirements:
Computer Science or related Engineering Degree (or above)
3 years or more working experience in SRE/DevOps duties
Central, Central and Western District, HK