About The Role
Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users across content, community, e-commerce, and advertising ecosystems. As the company expands internationally and deepens its investment in AI-driven infrastructure, observability is a critical backbone for platform reliability and performance. We are seeking a backend engineer with deep hands-on experience building observability systems at scale—not simply operating monitoring tools.
Location:
- Palo Alto, California
- 5 Days onsite
About The Role
Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users across content, community, e-commerce, and advertising ecosystems. As the company expands internationally and deepens its investment in AI-driven infrastructure, observability is a critical backbone for platform reliability and performance. We are seeking a backend engineer with deep hands-on experience building observability systems at scale—not simply operating monitoring tools.
You will help design and build next-generation observability infrastructure spanning metrics, logging, tracing, and profiling, supporting large-scale distributed systems as well as emerging AI infrastructure and AI-native workloads.
What You'll Work On
- Build and evolve large-scale observability systems across metrics, logging, tracing, and profiling.
- Design and develop monitoring platforms, distributed tracing systems, logging services, real-time alerting systems, and computation engines for streaming analytics and time-series workloads.
- Own architecture design and product-level implementation of core observability capabilities.
- Build systems designed for high throughput, high concurrency, low latency, high availability, and reliability at significant scale.
- Develop and improve service governance and observability capabilities across large-scale distributed and microservices environments.
- Work with cloud-native observability technologies such as OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF.
- Drive AI infrastructure observability, AI application observability, and 'Observability + AI' capabilities.
- Improve incident detection, troubleshooting, root-cause analysis, and overall platform stability through better telemetry and observability infrastructure.
What We're Looking For
- 2 -7 years of relevant backend engineering, infrastructure, or observability experience.
- Strong backend software engineering fundamentals and proficiency in Java or Go.
- Deep hands-on experience building observability, telemetry, monitoring, or reliability platforms—ideally at large scale.
- Strong understanding of distributed systems, concurrent programming, performance optimization, and high-concurrency system design.
- Practical experience with one or more of OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, or eBPF.
- Good understanding of Kubernetes and cloud-native infrastructure.
- Strong foundational knowledge of Linux, networking, storage systems, and message queues.
- Experience designing systems for high throughput, high availability, low latency, and large telemetry/data volumes.
- Strong coding ability, ownership, and ability to solve complex infrastructure problems end-to-end.
Language Requirement - Mandatory
- Fluent Mandarin Chinese is mandatory for this position.
- Candidates must be able to communicate effectively in Mandarin with engineering and product teams in China, while also being comfortable working in an English-speaking global engineering environment.
Preferred / Bonus Experience
- Experience building observability infrastructure for very large-scale consumer internet or distributed systems environments.
- Experience with AI infrastructure observability or AI application observability.
- Familiarity with AI-related technologies or ecosystems such as PyTorch, Spring AI, or Langfuse.
- Experience with cross-region or global infrastructure and international platform environments.
- Open-source contributions in observability, infrastructure, or distributed systems.
Location & Immigration:
- Palo Alto, California only.
- H-1B transfers can be supported for eligible US candidates.
- Green Card sponsorship/support can also be provided.
Why This Role
- Build foundational observability infrastructure supporting a global consumer platform with hundreds of millions of users.
- Work at the intersection of distributed systems, cloud-native observability, and AI infrastructure.
- Solve engineering problems involving massive telemetry volumes, high concurrency, performance, reliability, and global-scale infrastructure.
- Shape next-generation observability platforms rather than simply consuming existing monitoring tools.
- Gain direct exposure to AI observability—an emerging and highly specialized infrastructure domain.