M01 - Observability Engineer

FPT Asia Pacific

Singapore

On-site

SGD 90,000 - 120,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

FPT Asia Pacific is seeking an Observability Engineer to establish and operate consistent observability across SSOE infrastructure and applications. You will work across metrics, events, logs, and traces to provide a unified view of service health and performance.

You will define instrumentation standards, SLIs/SLOs, alerting strategies, dashboards, and operational health signals, collaborating with application, infrastructure, network, security, and platform teams to build observability in from

Qualifications

  • 3–5 years of experience in observability, SRE, or platform engineering.
  • Hands-on experience with metrics, logging, tracing, dashboards, and incident troubleshooting.
  • Experience working in on-premises and cloud environments with hybrid infrastructure.

Responsibilities

  • Design and operate end-to-end observability across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments.
  • Collect, aggregate, and correlate metrics, events, logs, and traces across infrastructure and application workloads.
  • Define and maintain observability standards across legacy, on-premise, containerised, and cloud-native systems.
  • Establish instrumentation standards using OpenTelemetry and other technologies.
  • Support engineering teams with instrumentation, SDK, agent, and telemetry integration.
  • Define service naming, metadata, tagging, correlation IDs, and telemetry enrichment.
  • Identify observability gaps and improve end-to-end visibility across services.
  • Define SLIs, SLOs, alerting rules, and service health indicators for critical services.
  • Build dashboards for infrastructure health, application performance, and reliability.
  • Develop leadership-level views of performance and trends.
  • Design actionable alerting to reduce noise while enabling rapid responses.
  • Establish monitoring requirements for new apps and platform components.
  • Use observability data for capacity planning, performance analysis, and operational decisions.
  • Define secure telemetry collection and routing across environments.

Skills

Observability engineering
SRE / Platform engineering
Metrics, logging, tracing
Incident troubleshooting
Hybrid cloud/on-prem
Cross-team collaboration

Tools

OpenTelemetry

Job description

Overview

As an Observability Engineer, you will establish and operate consistent observability capabilities across SSOE infrastructure and applications.

You will work across metrics, events, logs, and traces to provide a unified view of service health and performance. You will define instrumentation standards, service-level indicators and objectives, alerting strategies, dashboards, and operational health signals.

You will work closely with application, infrastructure, network, security, and platform teams to ensure observability is built into services from the outset rather than added after deployment.

Key Responsibilities
Observability Engineering
  • Design and operate end-to-end observability across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments
  • Collect, aggregate, and correlate metrics, events, logs, and traces across infrastructure and application workloads
  • Define and maintain observability standards that work consistently across legacy, on-premise, containerised, and cloud-native systems
  • Establish application and infrastructure instrumentation standards using OpenTelemetry and other appropriate technologies
  • Support engineering teams with instrumentation, SDK, agent, and telemetry integration
  • Define common conventions for service naming, metadata, tagging, correlation IDs, and telemetry enrichment
  • Identify observability gaps and continuously improve end-to-end visibility across SSOE services
Service Reliability & Monitoring
  • Define SLIs, SLOs, alerting rules, and service health indicators for critical services
  • Build operational dashboards covering infrastructure health, application performance, user experience, availability, capacity, and service reliability
  • Develop leadership-level views that provide meaningful visibility into service performance and operational trends
  • Design actionable alerting that enables teams to identify and respond to issues while minimising unnecessary alert noise
  • Establish monitoring and operational-readiness requirements for new applications, infrastructure, and platform components
  • Use observability data to support capacity planning, performance analysis, reliability improvements, and operational decision-making
Telemetry & Integration
  • Define secure telemetry collection and routing across on-premise environments, GCC, cloud platforms, and approved SaaS services
  • Work with infrastructure and platform teams to integrate telemetry from servers, network devices, applications, containers, databases, and managed cloud services
  • Design observability approaches that account for network boundaries, security zones, data residency, and connectivity constraints
  • Define telemetry retention, lifecycle, and cost-management requirements
  • Ensure logs, metrics, and traces can be correlated across distributed and hybrid systems
Incident Management & Continuous Improvement
  • Support operational teams during incidents by using observability data to identify symptoms, dependencies, and potential root causes
  • Participate in incident investigation, root-cause analysis, and post-incident reviews
  • Identify recurring operational issues and recommend improvements to instrumentation, alerting, architecture, or operational processes
  • Define appropriate SLOs and operational health indicators for observability services
  • Participate in operational support and on-call responsibilities for owned services
  • Maintain architecture documentation, operational procedures, and runbooks
What We Are Looking For
Experience
  • Minimum 3–5 years of experience in observability engineering, Site Reliability Engineering (SRE), platform engineering, infrastructure engineering, or a related discipline
  • At least 2 years of hands-on experience implementing or operating observability and monitoring capabilities in production environments
  • Demonstrated experience working with metrics, logging, tracing, dashboards, alerting, and incident troubleshooting
  • Experience monitoring and supporting production infrastructure, applications, or distributed systems
  • Experience working with on-premise and/or cloud environments, with an understanding of hybrid infrastructure
  • Experience working with engineering or operations teams to implement instrumentation and improve service reliability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

M01 - Observability Engineer
M01 - Observability Engineer

FPT ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Observability Engineer
Observability Engineer

U3 SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Software & Applications Manager (Technical Lead/Supervisor)
Software & Applications Manager (Technical Lead/Supervisor)

optimum solutions (singapore) pte ltd • Singapore

On-site
SGD 120,000 - 180,000
Observability Engineer
Observability Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Observability Solutions Architect
Observability Solutions Architect

Cisco • Singapore

On-site
SGD 80,000 - 120,000
Observability Engineer (Logging & Monitoring)
Observability Engineer (Logging & Monitoring)

Accenture Southeast Asia • Singapore

On-site
SGD 90,000 - 140,000
Observability Engineer: SRE‑Driven Monitoring & Reliability
Observability Engineer: SRE‑Driven Monitoring & Reliability

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 120,000
M01 - Cloud Logging & Data Platform Engineer
M01 - Cloud Logging & Data Platform Engineer

FPT ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Data & Observability Engineer for Hybrid Cloud Platforms
Data & Observability Engineer for Hybrid Cloud Platforms

Activate Interactive Pte Ltd. • Singapore

On-site
SGD 90,000 - 130,000
Competitive compensation
Flexible work arrangement
Learning & development opportunities
+2
M01 - Cloud Logging & Data Platform Engineer
M01 - Cloud Logging & Data Platform Engineer

FPT Asia Pacific • Singapore

On-site
SGD 120,000 - 180,000