Observability Backend Engineer - Distributed Systems (Mandarin required)

Applied Intelligence Consulting (Singapore)

San Francisco (CA)

On-site

USD 120,000 - 170,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intelligence Consulting (Singapore) is seeking a backend engineer to design and build observability infrastructure at scale for a global consumer internet platform. You will work on metrics, logging, tracing, and profiling across large distributed systems and AI workloads, ensuring high throughput and low latency.

You will implement using cloud-native tools like OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF, collaborating with cross-region teams.

Qualifications

  • 2 -7 years of relevant backend engineering, infrastructure, or observability experience.
  • Strong backend software engineering fundamentals and proficiency in Java or Go.
  • Deep hands-on experience building observability, telemetry, monitoring, or reliability platforms—ideally at large scale.
  • Strong understanding of distributed systems, concurrent programming, performance optimization, and high-concurrency system design.
  • Practical experience with one or more of OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, or eBPF.
  • Good understanding of Kubernetes and cloud-native infrastructure.
  • Strong foundational knowledge of Linux, networking, storage systems, and message queues.
  • Experience designing systems for high throughput, high availability, low latency, and large telemetry/data volumes.
  • Strong coding ability, ownership, and ability to solve complex infrastructure problems end-to-end.

Responsibilities

  • Build and evolve large-scale observability systems across metrics, logging, tracing, and profiling.
  • Design and develop monitoring platforms, distributed tracing systems, logging services, real-time alerting systems, and computation engines for streaming analytics and time-series workloads.
  • Own architecture design and product-level implementation of core observability capabilities.
  • Build systems designed for high throughput, high concurrency, low latency, high availability, and reliability at significant scale.
  • Develop and improve service governance and observability capabilities across large-scale distributed and microservices environments.
  • Work with cloud-native observability technologies such as OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF.
  • Drive AI infrastructure observability, AI application observability, and 'Observability + AI' capabilities.
  • Improve incident detection, troubleshooting, root-cause analysis, and overall platform stability through better telemetry and observability infrastructure.

Skills

Java
Go
Distributed systems
Observability platforms
OpenTelemetry
Prometheus
Kubernetes
Linux
Networking

Tools

OpenTelemetry
Prometheus
VictoriaMetrics
ELK
ClickHouse
SkyWalking
CAT
eBPF

Job description

About The Role

Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users across content, community, e-commerce, and advertising ecosystems. As the company expands internationally and deepens its investment in AI-driven infrastructure, observability is a critical backbone for platform reliability and performance. We are seeking a backend engineer with deep hands-on experience building observability systems at scale—not simply operating monitoring tools.

Location:
  • Palo Alto, California
  • 5 Days onsite
About The Role

Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users across content, community, e-commerce, and advertising ecosystems. As the company expands internationally and deepens its investment in AI-driven infrastructure, observability is a critical backbone for platform reliability and performance. We are seeking a backend engineer with deep hands-on experience building observability systems at scale—not simply operating monitoring tools.

You will help design and build next-generation observability infrastructure spanning metrics, logging, tracing, and profiling, supporting large-scale distributed systems as well as emerging AI infrastructure and AI-native workloads.

What You'll Work On
  • Build and evolve large-scale observability systems across metrics, logging, tracing, and profiling.
  • Design and develop monitoring platforms, distributed tracing systems, logging services, real-time alerting systems, and computation engines for streaming analytics and time-series workloads.
  • Own architecture design and product-level implementation of core observability capabilities.
  • Build systems designed for high throughput, high concurrency, low latency, high availability, and reliability at significant scale.
  • Develop and improve service governance and observability capabilities across large-scale distributed and microservices environments.
  • Work with cloud-native observability technologies such as OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF.
  • Drive AI infrastructure observability, AI application observability, and 'Observability + AI' capabilities.
  • Improve incident detection, troubleshooting, root-cause analysis, and overall platform stability through better telemetry and observability infrastructure.
What We're Looking For
  • 2 -7 years of relevant backend engineering, infrastructure, or observability experience.
  • Strong backend software engineering fundamentals and proficiency in Java or Go.
  • Deep hands-on experience building observability, telemetry, monitoring, or reliability platforms—ideally at large scale.
  • Strong understanding of distributed systems, concurrent programming, performance optimization, and high-concurrency system design.
  • Practical experience with one or more of OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, or eBPF.
  • Good understanding of Kubernetes and cloud-native infrastructure.
  • Strong foundational knowledge of Linux, networking, storage systems, and message queues.
  • Experience designing systems for high throughput, high availability, low latency, and large telemetry/data volumes.
  • Strong coding ability, ownership, and ability to solve complex infrastructure problems end-to-end.
Language Requirement - Mandatory
  • Fluent Mandarin Chinese is mandatory for this position.
  • Candidates must be able to communicate effectively in Mandarin with engineering and product teams in China, while also being comfortable working in an English-speaking global engineering environment.
Preferred / Bonus Experience
  • Experience building observability infrastructure for very large-scale consumer internet or distributed systems environments.
  • Experience with AI infrastructure observability or AI application observability.
  • Familiarity with AI-related technologies or ecosystems such as PyTorch, Spring AI, or Langfuse.
  • Experience with cross-region or global infrastructure and international platform environments.
  • Open-source contributions in observability, infrastructure, or distributed systems.
Location & Immigration:
  • Palo Alto, California only.
  • H-1B transfers can be supported for eligible US candidates.
  • Green Card sponsorship/support can also be provided.
Why This Role
  • Build foundational observability infrastructure supporting a global consumer platform with hundreds of millions of users.
  • Work at the intersection of distributed systems, cloud-native observability, and AI infrastructure.
  • Solve engineering problems involving massive telemetry volumes, high concurrency, performance, reliability, and global-scale infrastructure.
  • Shape next-generation observability platforms rather than simply consuming existing monitoring tools.
  • Gain direct exposure to AI observability—an emerging and highly specialized infrastructure domain.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Observability Backend Engineer for AI-Driven Infra
Observability Backend Engineer for AI-Driven Infra

Applied Intelligence Consulting (Singapore) • San Francisco (CA)

On-site
USD 120,000 - 170,000
Backend Engineer
Backend Engineer

Space Executive • Berkeley (CA)

Remote
USD 125,000 - 225,000
Medical, dental, vision
401(k)
Unlimited PTO
+2
Observability Engineer - Telemetry Extension & ADOT Pipeline
Observability Engineer - Telemetry Extension & ADOT Pipeline

Intellias • Spain (TX)

On-site
EUR 70,000 - 120,000
Software Engineer, Observability
Software Engineer, Observability

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 175,000
Equity
Healthcare
Mentorship & events
+2
Principal Platform Engineer, Observability (CIPE)
Principal Platform Engineer, Observability (CIPE)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 147,000 - 238,000
Backend Engineer, Distributed Systems (English/Mandarin - Fully Bilingual requirement)
Backend Engineer, Distributed Systems (English/Mandarin - Fully Bilingual requirement)

Applied Intelligence Consulting (Singapore) • San Francisco (CA)

On-site
USD 180,000 - 240,000
Principal Full Stack Software Engineer (Observability)
Principal Full Stack Software Engineer (Observability)

Salesforce • San Francisco (CA)

On-site
USD 180,000 - 260,000
Medical Care
Life Insurance
Retirement Savings
+2
Engineering Manager, Observability Platforms
Engineering Manager, Observability Platforms

TEKsystems • Raleigh (NC)

On-site
USD 169,000 - 200,000
401(k) match up to 5%
18 days PTO with option to buy an add.
Flexible hours
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000