Senior Observability Engineer (AI GPU Cloud) - Telemetry/Prometheus/Cutting-edge technology

Dada Consultants

Singapore

On-site

SGD 120,000 - 180,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Dada Consultants is seeking a Senior Observability Engineer for a leading AI cloud infrastructure provider in Singapore. You will own the telemetry stack from hardware metrics collection to multi-tenant billing pipelines, shaping throughput and reliability.

Responsibilities include building scalable observability, integrating kernel-level tracing, and developing dashboards and alerting to detect degraded hardware before it impacts workloads.

Qualifications

  • Bachelor's or master's degree in CS/EE or a related discipline.
  • Strong experience with Prometheus/OpenTelemetry and Go in large-scale environments.
  • Experience with Kubernetes metric exporters and operators, kernel tracing tools (eBPF/BCC).

Responsibilities

  • Architect and scale a high-throughput observability platform across a large distributed compute environment.
  • Integrate hardware telemetry (accelerator health, network fabric statistics) into a cluster-wide monitoring layer.
  • Build kernel-level diagnostic tooling to trace network congestion, I/O latency, and bottlenecks.
  • Develop dashboards and alerting pipelines to identify degraded hardware before workloads are affected.
  • Design metering pipelines for real-time, multi-tenant usage-based billing.
  • Partner with infra and scheduling teams to define observability standards for AI/ML workloads.
  • Lead architecture reviews and mentor engineers on high-performance telemetry collection.

Skills

Prometheus
OpenTelemetry
Go
Kubernetes
eBPF/BCC
Linux performance
AI hardware metrics

Education

Bachelor's or Master's degree in Computer Science or Electrical Engineering

Tools

Kubernetes exporters
Kernel tracing tools
Linux performance tuning

Job description

Our client is a large-scale AI cloud infrastructure provider that offers GPU compute capacity, high-performance training and inference infrastructure. As a Senior Observability Engineer, you will own the design and scaling of the company's entire telemetry stack — from hardware-level metrics collection through to multi-tenant billing pipelines.

Key Responsibilities
  • Architect and scale a high-throughput observability platform capable of ingesting and querying extremely high-cardinality metrics reliably across a large, distributed compute environment.
  • Integrate low-level hardware telemetry (accelerator health, network fabric statistics, and out-of-band system management data) into a unified cluster-wide monitoring layer.
  • Build custom kernel-level diagnostic tooling to trace network congestion, I/O latency, and distributed workload bottlenecks across large compute clusters.
  • Develop automated dashboards and alerting pipelines that proactively identify and isolate degraded hardware before it impacts running workloads.
  • Design metering pipelines that support accurate, multi-tenant usage-based billing derived from real-time compute and network utilization data.
  • Partner with infrastructure and scheduling teams to define observability standards for large-scale AI/ML workloads, ensuring visibility into job-level efficiency and resource usage.
  • Lead architecture and design reviews for the observability stack, mentoring engineers on best practices for high-performance telemetry collection and analysis.
Requirements
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related discipline
  • Software or site reliability engineering experience, with expertise in Prometheus/OpenTelemetry
  • Advanced proficiency in Go, with substantial experience in Kubernetes metric exporters and operators
  • Familiarity with kernel-level tracing tools (eBPF, BCC) and performance tuning of Linux systems at scale
  • Strong familiarity with AI hardware performance metrics (GPU power states, compute utilisation, memory bandwidth) and high-performance network telemetry

EA License No.: 18S9037

Business Registration Number: 201735941W

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Telemetry & Observability Engineer
Senior AI Telemetry & Observability Engineer

Bitdeer • Singapore

On-site
SGD 120,000 - 180,000
Senior Observability Engineer for AI Infra Telemetry
Senior Observability Engineer for AI Infra Telemetry

Dada Consultants • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Observability & Telemetry Engineer
Senior AI Observability & Telemetry Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Senior AI Infra Engineer, Observability
Senior AI Infra Engineer, Observability

Firmus Technologies • Singapore

On-site
SGD 180,000 - 240,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Infrastructure Engineer, Observability
Senior AI Infrastructure Engineer, Observability

Firmus • Singapore

On-site
SGD 180,000 - 260,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
6723 - AI Infrastructure Engineer
6723 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI DevOps Engineer (Cloud Infrastucture)
AI DevOps Engineer (Cloud Infrastucture)

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000