Lead Data Engineer

LionsBot International

Singapore

On-site

SGD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LionsBot International is building autonomous cleaning robots for malls, airports, offices and industrial sites. We seek a data platform owner to take telemetry streams and turn them into a trusted, streaming-first data platform with a clear medallion architecture and semantic layer.

You will own ingestions to dashboards across 30+ countries. We expect 3+ years in production data roles, strong SQL and time-series design, and experience with real-time data pipelines.

Qualifications

  • 3+ years working with data in production: data engineering, analytics engineering, or backend with heavy data exposure.
  • Strong SQL and PostgreSQL with scalable data models and time-series focus.
  • Experience with streaming or real-time data pipelines and dashboards.
  • Ability to ship dashboards and metrics that teams actually use.

Responsibilities

  • Own the data platform end-to-end: ingestion, storage, modeling, serving, dashboards.
  • Design medallion architecture (bronze, silver, gold) for IoT telemetry.
  • Build metrics/semantic layer with clearly defined KPIs.
  • Run data quality, observability, and anomaly alerts like production software.
  • Architect databases at scale: schemas, indexes, aggregates, retention.
  • Make analytics self-serve with dashboards for ops, product, leadership.
  • Shape the roadmap for OLAP, orchestration, lakehouse patterns.
  • Collaborate with AI agents to ensure safe, correct querying.

Skills

SQL
PostgreSQL
Time-series DB
Kafka
Python
Go/Java
Dashboards

Tools

TimescaleDB
InfluxDB
Elasticsearch

Job description

Your data sources have wheels. LionsBot designs and builds autonomous cleaning robots that work in the real world: malls, airports, offices and industrial sites across 30+ countries. More than 5,000 of them stream telemetry to us in real time: missions, maps, locations, incidents, battery health. We move fast: small team, quick decisions, zero bureaucracy, hardware you can kick.

You'll be our first dedicated data hire, a true 0→1, greenfield ownership role. The foundations are in place: real-time telemetry streams from the fleet into a time-series store, with dashboards on top. Our fleet has now grown to the point where data deserves a full-time owner, so we're making it a first-class function. End-to-end, it's yours.

The mission: take us from "a pipeline that works" to a streaming-first data platform with a proper medallion architecture: bronze raw telemetry, silver cleaned and conformed, gold business-ready marts. On top of it all, a semantic layer where every metric has exactly one definition and everyone trusts the number.

The fun problems, all real

  • Robots report cumulative lifetime odometers on every mission row. Sum the wrong column and your fleet total inflates 1,000×. Design the models that make that mistake impossible.
  • A sensor glitch claims one robot cleaned 2.5 million m² in twenty minutes. Build the data quality and anomaly detection that catches it before a human ever sees it.
  • Robots in basements with bad Wi-Fi send late-arriving, out-of-order data. Make the pipelines idempotent anyway.
  • Real-time fleet health: which robots are sick right now, across 30+ countries and time zones?
What you will do
  • Own the data platform end-to-end: ingestion, storage, modeling, serving, dashboards. Real-time event streams from the fleet land in a time-series database today. Where it goes next is your call.
  • Design the medallion architecture: bronze, silver and gold layers over high-volume IoT telemetry, with clear data contracts agreed with the backend teams so quality is designed in at the source.
  • Build the metrics/semantic layer: canonical, documented, version-controlled definitions for fleet KPIs: cleaning hours, area, mission success, incident rates, robot health. A genuine single source of truth.
  • Run data quality & observability like production software: freshness SLAs, validation, dedup, outlier handling, anomaly alerts. Flag the weird number before leadership does.
  • Design, tune and re-architect databases at scale: schemas, indexes, continuous aggregates, compression, downsampling, retention and partitioning, treating them like the production systems they are.
  • Make analytics self-serve: dashboards and models for ops, product, leadership and OEM partners, plus fast, rigorous answers to the high-stakes ad-hoc questions.
  • Shape the roadmap: we run lean today, so what comes next is genuinely open: OLAP, orchestration, transformation tooling, lakehouse patterns. You evaluate, make the case, and we adopt what earns its keep.
  • Work AI-native: we pair humans with LLM-powered analytics agents daily. You'll design the platform so both humans and AI agents can query it safely and correctly.
What we are looking for
  • 3+ years working with data in production: data engineering, analytics engineering, or backend with heavy data exposure. We hire for trajectory, not year count.
  • Strong SQL, solid PostgreSQL and confident database design: schemas, indexes and data models that hold up as data grows. Time-series databases like TimescaleDB or InfluxDB are a big plus, but you'll learn them fast here.
  • Comfortable with event-driven data: you've worked with streaming or message-queue systems like Kafka, or you're a data-minded backend engineer keen to go deeper on real-time.
  • Solid Python for pipelines and tooling, and comfortable reading Go or Java services.
  • You've shipped dashboards and metrics people actually used, whatever the BI tool.
  • Fast and autonomous, like our robots: high ownership, pragmatic trade-offs, comfortable with ambiguity, ships iteratively.
Nice to have:
  • Production streaming chops: you know your at-least-once from your exactly-once.
  • IoT, robotics, or high-volume device telemetry experience.
  • AWS, especially EKS, RDS and S3, with exposure to Azure or GCP.
  • OLAP engines, orchestration or transformation tooling, CDC pipelines.
  • Search engines like Quickwit or Elasticsearch, or graph databases like Neo4j.
  • Geospatial data: our robots navigate real floors, so maps and location streams are first-class citizens.
  • Experience making data platforms LLM/agent-friendly: semantic layers, governed self-serve.
  • Familiarity with OpenRMF, ROS or robotics-related communication stacks
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

Neara • Singapore

On-site
SGD 120,000 - 180,000
On-site work
Lead Data Platform Engineer - Real-Time IoT & Robotics
Lead Data Platform Engineer - Real-Time IoT & Robotics

LionsBot International • Singapore

On-site
SGD 120,000 - 190,000
Senior Full Stack Engineer
Senior Full Stack Engineer

Hydrax • Singapore

On-site
SGD 85,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Hyundai Motor Group Innovation Center Singapore (HMGICS) • Singapore

On-site
SGD 120,000 - 180,000
Analytics Engineer
Analytics Engineer

Workato • Singapore

On-site
SGD 80,000 - 120,000
AI Data Platform Lead
AI Data Platform Lead

NEWBRIDGE ALLIANCE PTE. LTD. • Singapore

On-site
SGD 180,000 - 300,000
Senior Data Engineer
Senior Data Engineer

michael page (personnel) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Data Engineering & AI - FinTech Investments / Unstructured Data
Data Engineering & AI - FinTech Investments / Unstructured Data

Gravitas Recruitment Group (Global) Ltd • Singapore

On-site
SGD 120,000 - 150,000
Senior Software Engineer
Senior Software Engineer

the trade desk (singapore) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)

Xora Innovation • Singapore

On-site
SGD 180,000 - 280,000