Get more replies from employers
Send a job-specific resume in minutes.
NEWBRIDGE ALLIANCE PTE. LTD. in Singapore is seeking a Staff / Lead Data Engineer (AI-Native) to architect petabyte-scale data pipelines feeding real-time bidding, recommendations, and analytics.
You will balance hands-on coding with mentoring a small team, focusing on reliability, scalability, and AI-ready data foundations built on Spark, Flink, Databricks, and modern lakehouse architectures.
Our client is a high-scale global consumer platform serving millions of users daily. They are rebuilding their data foundation as an AI-native platform to power the next generation of programmatic advertising, personalized discovery, and marketplace intelligence.
We are hiring a hands-on Staff / Lead Data Engineer (AI-Native) to architect petabyte-scale systems that feed real-time bidding, recommendation models, and LLM-powered analytics products.
This is NOT a pure management role. You will be a player-coach - designing and shipping critical pipelines yourself while mentoring a small pod of engineers. Strong individual contributors with no formal people management experience but with deep technical depth are strongly encouraged to apply.
Design, build, and operate AI-ready batch + streaming ETL/ELT pipelines ingesting 100TB+ daily from ad servers, mobile SDKs, transactional systems, and 3rd-party APIs. Build for LLM and ML consumption from day one.
Develop low-latency streaming jobs using Spark Streaming, Flink, or Kafka Streams for real-time use cases: fraud detection, bid optimization, dynamic pricing, and real-time personalization. Enable online inference and agentic decisioning.
Model and optimize massive datasets on a modern lakehouse [Databricks / Snowflake / BigQuery] to serve BI, embedded analytics, and AI/ML workloads with a focus on performance, cost, and feature freshness.
Build reliable data products for Data Science & ML: feature stores [Feast / Tecton], vector stores for semantic search & RAG, training datasets with point-in-time correctness, and online-offline parity.
Own data quality, observability, anomaly detection, and lineage for Tier-0 datasets that directly power revenue, user experience, and model performance.