Observability Engineer (Site Reliability Engineering)

Horizontal Talent

Petaling Jaya

On-site

MYR 140,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

HAVI is seeking an Observability Engineer to build and evolve enterprise-wide observability capabilities across platforms and services in a cloud-native environment. You will make system health, performance and reliability visible, measurable and actionable by developing logging, metrics, tracing and alerting capabilities.

You will design instrumentation standards, build dashboards aligned with SLOs, optimise alerting, and collaborate with SRE and development teams to embed observability into

Qualifications

  • 4+ years of experience in Monitoring, Observability or SRE-related roles.
  • Strong experience with observability platforms such as Azure Monitor, Prometheus, ELK, Datadog and Splunk.
  • Strong understanding of metrics, logs and distributed tracing.
  • Experience designing scalable telemetry architectures.
  • Strong knowledge of alert design and signal-to-noise optimisation.
  • Familiarity with cloud-native monitoring integrations.
  • Scripting experience with Python, Bash or similar.
  • Understanding of SLO-driven monitoring strategies.
  • Strong analytical and data interpretation skills.
  • Bachelor's degree in Computer Science, Engineering or equivalent experience.

Responsibilities

  • Design and maintain enterprise logging, metrics and tracing platforms
  • Define instrumentation standards for applications and infrastructure
  • Build dashboards aligned with SLOs and service health indicators
  • Optimise alerting frameworks to improve signal quality and reduce noise
  • Ensure telemetry pipelines are scalable, reliable and cost-efficient
  • Partner with SRE teams to improve visibility into error budgets and performance trends
  • Work with Application Development teams to embed observability into system design
  • Continuously improve detection speed and diagnostic capabilities

Skills

Observability
SRE
Monitoring
Python

Education

Bachelor's degree in Computer Science, Engineering or equivalent

Tools

Azure Monitor
Prometheus
ELK
Datadog
Splunk

Job description

About The Role

HAVI is looking for an Observability Engineer to help build and evolve enterprise-wide observability capabilities across platforms and services. You'll be responsible for making system health, performance and reliability visible, measurable and actionable by building and maintaining logging, metrics, tracing and alerting capabilities. This is a great opportunity for someone with a strong background in Observability, Monitoring or SRE who enjoys improving system visibility and reliability in cloud-native and distributed environments.

What You'll Do
  • Design and maintain enterprise logging, metrics and tracing platforms
  • Define instrumentation standards for applications and infrastructure
  • Build dashboards aligned with SLOs and service health indicators
  • Optimise alerting frameworks to improve signal quality and reduce noise
  • Ensure telemetry pipelines are scalable, reliable and cost-efficient
  • Partner with SRE teams to improve visibility into error budgets and performance trends
  • Work with Application Development teams to embed observability into system design
  • Continuously improve detection speed and diagnostic capabilities
What We're Looking For
  • 4+ years of experience in Monitoring, Observability or SRE-related roles
  • Strong experience with observability platforms such as:
    • Azure Monitor
    • Prometheus
    • ELK
    • Datadog
    • Splunk
  • Strong understanding of metrics, logs and distributed tracing
  • Experience designing scalable telemetry architectures
  • Strong knowledge of alert design and signal-to-noise optimisation
  • Familiarity with cloud-native monitoring integrations
  • Scripting experience with Python, Bash or similar
  • Understanding of SLO-driven monitoring strategies
  • Strong analytical and data interpretation skills
  • Bachelor's degree in Computer Science, Engineering or equivalent experience
Why Consider This Role?

You'll play a key role in building HAVI's enterprise observability capability, working with SRE, Platform Engineering, Application Development, Security and Service Operations teams in a global environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Engineer: Enterprise Telemetry
SRE Observability Engineer: Enterprise Telemetry

Horizontal Talent • Petaling Jaya

On-site
MYR 140,000 - 180,000
SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
SRE Lead
SRE Lead

Chubb Ltd. • Malaysia

On-site
MYR 240,000 - 420,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Hong Kong and Macau • Kuala Lumpur

On-site
MYR 70,000 - 90,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
Lead Site Reliability Engineer ELK
Lead Site Reliability Engineer ELK

Horizontal Talent • Kuala Lumpur

On-site
MYR 240,000 - 420,000
Observability Solutions Architect
Observability Solutions Architect

Cisco • Kuala Lumpur

On-site
MYR 180,000 - 320,000
Observability Solutions Architect
Observability Solutions Architect

Cisco Systems, Inc. • Kuala Lumpur

Hybrid
MYR 350,000 - 520,000
SRE Lead
SRE Lead

Chubb • Malaysia

On-site
MYR 300,000 - 420,000
Site Reliability Engineer (SRE) - Ads / Monetization Platform
Site Reliability Engineer (SRE) - Ads / Monetization Platform

Two95 International Inc. • Kuala Lumpur

On-site
MYR 120,000 - 180,000