Primary Purposes::
We are looking for an experienced Data Engineer to join the delivery team for the EDW Reporting Modernisation project. You will be responsible for designing and building the Gold layer Data Mart tables, transformation pipelines, and associated views on the Unified Data Platform (Databricks), working closely with the Tech Lead and Data Modeler to deliver a robust and scalable reporting data layer that caters to Power BI reports.
As the onshore Data Engineer, you will serve as the technical bridge between the Singapore-based design team and the offshore engineering squad, providing day-to-day technical direction, code review, and quality oversight across the full Data Mart build.
Responsibilities::
Design and Development::
- Develop Gold layer Data Mart tables and associated database views in Databricks in accordance with signed-off Source-to-Target Mapping (STTM) documents and design specifications
- Implement Silver-to-Gold transformation pipelines using Databricks Workflows, Delta Live Tables (DLT), and PySpark
- Create Delta table DDL in Unity Catalog, including schema definition, table properties, and column-level comments
- Implement SCD Type 2 merge logic for dimension tables and aggregation logic for fact tables per the grain specification
- Configure table partitioning, Z-ORDER clustering, and Liquid Clustering strategies for query performance optimisation
- Develop and implement data quality expectations using Delta Live Tables DQ framework per the DQ thresholds defined in the design specification
- Schedule and configure Databricks Workflow job DAGs for all pipeline runs, including watermark-based incremental load logic
Testing and Quality Assurance::
- Conduct unit testing for each Data Mart table upon build completion — covering row count validation, null checks, duplicate grain checks, and transformation logic verification
- Execute system integration test (SIT) scenarios for all 60 Data Mart tables, including aggregated value checks, referential integrity validation, and SCD integrity checks
- Log, investigate, and resolve defects identified during SIT and UAT, performing root cause analysis and documenting fixes
- Review and validate unit test outputs from Pune Data Engineers before test results are submitted for Tech Lead review
Offshore Squad Technical Leadership::
- Provide day-to-day technical direction and coding guidance to the 2 offshore-based Data Engineers
- Conduct code reviews for all pipeline and DDL code produced by the offshore squad before submission for Tech Lead sign-off
- Ensure offshore squad adherence to agreed coding standards, naming conventions, Unity Catalog governance rules, and development best practices
- Facilitate knowledge sharing and technical problem-solving with the Pune squad on complex transformation and merge logic
Deployment and Handover::
- Execute Data Mart pipeline and table deployment to the production Databricks environment per the approved deployment runbook
- Validate the first production pipeline run and confirm data freshness post-deployment
- Contribute to the operational runbook covering pipeline architecture, job schedules, alert thresholds, and common failure scenarios
- Support knowledge transfer sessions with the Singtel IT operations team during the handover phase
Qualifications
Required Skills and Experience
Technical — Mandatory
- 5+ years of experience in data engineering with a strong focus on ETL/ELT pipeline development and dimensional data modelling
- Hands-on experience with Databricks — including Delta Lake, Delta Live Tables, Databricks Workflows, Unity Catalog, and Databricks SQL
- Proficiency in PySpark and Databricks SQL for large-scale data transformation
- Strong understanding of dimensional modelling concepts — star schema, SCD Type 2, surrogate key design, fact and dimension table design
- Experience implementing data quality frameworks and reconciliation testing
- Familiarity with Source-to-Target Mapping (STTM) and translating design specifications into production-grade pipelines
- Experience with incremental load patterns — watermark-based, partition-based, or CDC-driven
- Proficiency in Git-based version control for collaborative development
Technical — Preferred::
- Experience with Databricks Unity Catalog access control, table tagging, and lineage tracking
- Exposure to Oracle-to-Databricks migration projects or similar platform modernisation programmes
- Familiarity with Power BI Semantic Models and how Gold layer table design impacts downstream DAX measure performance
- Experience with Databricks Assistant or AI-assisted code generation tools for accelerated pipeline development
- Knowledge of CI/CD pipeline setup for Databricks notebook or YAML