Observability Backend Engineer for AI-Driven Infra

Applied Intelligence Consulting (Singapore)

San Francisco (CA)

On-site

USD 120,000 - 170,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intelligence Consulting (Singapore) is seeking a backend engineer to design and build observability infrastructure at scale for a global consumer internet platform. You will work on metrics, logging, tracing, and profiling across large distributed systems and AI workloads, ensuring high throughput and low latency.

You will implement using cloud-native tools like OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF, collaborating with cross-region teams.

Qualifications

  • 2 -7 years of relevant backend engineering, infrastructure, or observability experience.
  • Strong backend software engineering fundamentals and proficiency in Java or Go.
  • Deep hands-on experience building observability, telemetry, monitoring, or reliability platforms—ideally at large scale.
  • Strong understanding of distributed systems, concurrent programming, performance optimization, and high-concurrency system design.
  • Practical experience with one or more of OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, or eBPF.
  • Good understanding of Kubernetes and cloud-native infrastructure.
  • Strong foundational knowledge of Linux, networking, storage systems, and message queues.
  • Experience designing systems for high throughput, high availability, low latency, and large telemetry/data volumes.
  • Strong coding ability, ownership, and ability to solve complex infrastructure problems end-to-end.

Responsibilities

  • Build and evolve large-scale observability systems across metrics, logging, tracing, and profiling.
  • Design and develop monitoring platforms, distributed tracing systems, logging services, real-time alerting systems, and computation engines for streaming analytics and time-series workloads.
  • Own architecture design and product-level implementation of core observability capabilities.
  • Build systems designed for high throughput, high concurrency, low latency, high availability, and reliability at significant scale.
  • Develop and improve service governance and observability capabilities across large-scale distributed and microservices environments.
  • Work with cloud-native observability technologies such as OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF.
  • Drive AI infrastructure observability, AI application observability, and 'Observability + AI' capabilities.
  • Improve incident detection, troubleshooting, root-cause analysis, and overall platform stability through better telemetry and observability infrastructure.

Skills

Java
Go
Distributed systems
Observability platforms
OpenTelemetry
Prometheus
Kubernetes
Linux
Networking

Tools

OpenTelemetry
Prometheus
VictoriaMetrics
ELK
ClickHouse
SkyWalking
CAT
eBPF

Job description

Applied Intelligence Consulting (Singapore) is seeking a backend engineer to design and build observability infrastructure at scale for a global consumer internet platform. You will work on metrics, logging, tracing, and profiling across large distributed systems and AI workloads, ensuring high throughput and low latency.

You will implement using cloud-native tools like OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF, collaborating with cross-region teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Backend Engineer - Distributed Systems (Mandarin required)
Observability Backend Engineer - Distributed Systems (Mandarin required)

Applied Intelligence Consulting (Singapore) • San Francisco (CA)

On-site
USD 120,000 - 170,000
Backend Engineer for AI Observability Platform
Backend Engineer for AI Observability Platform

AnoSys Technologies, Inc • Los Angeles (CA)

On-site
USD 100,000 - 140,000
Senior Observability Engineer - AI Telemetry Platform
Senior Observability Engineer - AI Telemetry Platform

CoreWeave • Sunnyvale (CA)

Hybrid
USD 165,000 - 242,000
Senior Site Reliability Engineer - AI Observability & Infra
Senior Site Reliability Engineer - AI Observability & Infra

Tulip Interfaces • Boston (MA)

Hybrid
USD 160,000 - 200,000
Company equity
Flexible schedule
Health, dental, vision
Observability Engineer: Scale Telemetry & AI Reliability
Observability Engineer: Scale Telemetry & AI Reliability

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Competitive compensation including 1x/
Full medical/dental/vision insurance
Flexible PTO with Winter Break
+3
Senior Observability Lead for AI-Driven Infra
Senior Observability Lead for AI-Driven Infra

Tulip Interfaces • Boston (MA)

Hybrid
USD 160,000 - 200,000
Health insurance
Dental coverage
Vision coverage
+12
Observability Engineer - Telemetry Extension & ADOT Pipeline
Observability Engineer - Telemetry Extension & ADOT Pipeline

Intellias • Spain (TX)

On-site
EUR 70,000 - 120,000
Staff Observability Engineer - Scale, AI-Driven Systems
Staff Observability Engineer - Scale, AI-Driven Systems

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Senior Observability Architect & AI-Ops Lead
Senior Observability Architect & AI-Ops Lead

Jobtailor • Town of Florida (NY)

On-site
USD 150,000 - 210,000
Staff Product Manager: AI Agent & Systems Observability
Staff Product Manager: AI Agent & Systems Observability

Webhosting • United States

On-site
USD 150,000 - 210,000