Senior Data Engineer - Data Lakehouse

ONL Biz Solutions Sdn Bhd

Kuala Lumpur

On-site

MYR 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ONL Biz Solutions Sdn Bhd is seeking a senior data engineer to own end-to-end data pipelines and governance for a lakehouse platform. You will work with Iceberg tables on AWS, Spark/Flink processing, and a ClickHouse serving layer to enable fast analytics.

Ideal candidates have 6+ years in data engineering, strong SQL and data modeling, and hands-on experience with distributed compute. Excellent collaboration and communication are essential for cross-team work.

Qualifications

  • 6+ years in data engineering or data-intensive backend roles.
  • Production experience with a lakehouse/open table format.
  • Strong distributed compute: Spark and/or Flink; performance tuning a plus.
  • Expert SQL and dimensional data modeling; proficient with analytical stores.
  • Solid AWS fundamentals: S3, IAM, VPC, RDS.
  • Production reliability mindset: monitoring, alerts, on-call, root-cause analysis.
  • Proficiency in Python, Java, or Scala for pipelines.
  • Excellent communication and stakeholder-management skills.

Responsibilities

  • Own and operate the streaming platform for reliability, correctness, and cost.
  • Build and tune batch (Spark) and real-time (Flink) processing pipelines.
  • Develop and optimize the serving layer (ClickHouse over Iceberg) for fast queries.
  • Design data models, partitioning, and compaction to balance write and query costs.
  • Establish data governance: catalog, access control, data lineage, retention, quality checks.
  • Collaborate with analysts, product teams, and engineers for accessible data access.
  • Stay current with lakehouse ecosystem and propose improvements.

Skills

SQL
Python
Java
Scala
Spark
Flink
ClickHouse
AWS
Data modeling
Data governance

Tools

Iceberg
Kafka Connect
Flink
Spark
ClickHouse
S3
Terraform
Kubernetes
Dagster
Airflow

Job description

About the role

This is a hands-on, senior individual-contributor role with a strong passion on building a possible vendor neutral lakehouse. You will own data pipelines end to end and build out the governance and data-quality layer that a fast-growing platform now needs.

Our stack (so you know what you'd be working with)

Lakehouse: Apache Iceberg tables on S3, Apache Polaris (Iceberg REST catalog); Kafka Connect and Apache Flink write into it.

Processing: Apache Flink (real-time upserts / streaming) and Apache Spark (batch silver/gold), both reading and writing Iceberg. Serving / query: ClickHouse

Platform / IaC: AWS, Amazon EKS (Kubernetes), Terraform, Helm; observability via Prometheus, Grafana, and Alertmanager.

We value experience with equivalent tools too — if you've done Delta/Hudi instead of Iceberg, Kinesis instead of Kafka, or Airflow instead of Dagster, that transfers.

Key responsibilities

Own, operate and harden the streaming platform for reliability, correctness, and cost at scale.

Build and tune the processing layers: batch transformations in Spark (silver/gold models) and real-time streaming jobs in Flink, including partitioning, compaction, and job/state tuning.

Develop and optimize the serving layer (ClickHouse over Iceberg) so analysts and product teams get fast, reliable query access.

Design data models and physical table layout (partitioning, sort order, compaction, retention) to balance write throughput, query performance, and storage cost across the lakehouse.

Establish and enforce data governance: catalog organization, access control, data lineage, retention/compliance, and automated data-quality checks.

Collaborate with data analysts, product teams, and backend engineers to ensure discoverable, well-documented, and usable data access.

Stay current with the lakehouse ecosystem and propose targeted improvements.

Requirements

6+ years in data engineering or backend data-intensive roles.

Production experience with an open table format / lakehouse.

Strong distributed compute: Apache Spark and/or Apache Flink for batch and/or stream processing, including performance tuning is a plus.

Expert SQL, dimensional/data modeling, and query performance tuning — including experience with an analytical/columnar store (e.g., ClickHouse, Redshift, Snowflake, BigQuery).

Solid AWS fundamentals: S3, IAM, VPC/networking, and RDS.

A production-reliability mindset: monitoring, alerting, on-call, and incident root-cause analysis.

Proficiency in at least one of Python, Java, or Scala for pipeline and streaming code.

Excellent communication and stakeholder-management skills.

Nice to have

Real-time OLAP experience (ClickHouse, Apache Druid, Apache Pinot).

Workflow orchestration (Dagster, Airflow, or equivalent).

Multi-region and/or cross-account data platform experience.

Data governance, lineage, and quality tooling (OpenLineage/Marquez, Great Expectations, dbt tests, catalog-based access control).

Security and compliance frameworks (GDPR, HIPAA, SOC 2, etc.).

Observability stacks (Prometheus, Grafana, Alertmanager).

Cost optimization for cloud data/streaming platforms.

ONL Biz Solutions Sdn Bhd is a young and innovative IT software warehouse company based in Kuala Lumpur. We are driven by a vision to empower the younger generation in the technology era. Our team consists of passionate programmers and designers who thrive on new challenges and constantly seek innovation.

We pride ourselves on being a dynamic and rapidly expanding company. As part of our mission to grow, we are actively seeking talented individuals to join our team in Kuala Lumpur. We look forward to meeting you and exploring potential opportunities.

ONL Biz Solutions Sdn Bhd is a young and innovative IT software warehouse company based in Kuala Lumpur. We are driven by a vision to empower the younger generation in the technology era. Our team consists of passionate programmers and designers who thrive on new challenges and constantly seek innovation.

We pride ourselves on being a dynamic and rapidly expanding company. As part of our mission to grow, we are actively seeking talented individuals to join our team in Kuala Lumpur. We look forward to meeting you and exploring potential opportunities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer — End-to-End Lakehouse Platform
Senior Data Engineer — End-to-End Lakehouse Platform

ONL Biz Solutions Sdn Bhd • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Senior Data Engineer
Senior Data Engineer

MetaComp • Kuala Lumpur

On-site
MYR 479,000 - 607,000
Data Engineer
Data Engineer

Royal Bank of Canada • Putrajaya

On-site
MYR 120,000 - 240,000
Senior Data Engineer: Lakehouse & Data Warehouse Lead
Senior Data Engineer: Lakehouse & Data Warehouse Lead

KLK OLEO • Selangor

On-site
MYR 120,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Tranglo • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Senior Data Platform Engineer – Lakehouse & Analytics
Senior Data Platform Engineer – Lakehouse & Analytics

MetaComp • Kuala Lumpur

On-site
MYR 479,000 - 607,000
Data Engineer
Data Engineer

Oxydata Software Sdn Bhd • Kuala Lumpur

On-site
Data Engineers (Junior to Lead Levels)
Data Engineers (Junior to Lead Levels)

Randstad Malaysia • Kuala Lumpur

On-site
Health insurance
Performance bonuses
Continuous learning allowances
Data Engineer
Data Engineer

Oxydata Software • Petaling Jaya

On-site
MYR 90,000 - 150,000
Competitive salary
Onsite team in Kuala Lumpur
Assistant Manager-Senior Executive Data Engineering
Assistant Manager-Senior Executive Data Engineering

QL Resources Berhad • Shah Alam

On-site
MYR 120,000 - 180,000