- Production delivery of data pipelines inside banks
- Customer, account, transaction, payments and AML data
- Working within bank release, scheduling and change controls
- Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning
- Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R
- Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports
- ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation
- Reconciliation, data quality and SLA monitoring
- Performance at scale: tables of 1 billion+ rows and multi-year history
Role: Lead Data Engineer (Banking)
Must-have Skills
Banking and domain
- Production delivery of data pipelines inside banks
- Customer, account, transaction, payments and AML data
- Working within bank release, scheduling and change controls
- Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning
- Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R
- Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports
- ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation
- Reconciliation, data quality and SLA monitoring
- Performance at scale: tables of 1 billion+ rows and multi-year history
Core data engineering
- Production delivery of data pipelines inside banks
- Customer, account, transaction, payments and AML data
- Working within bank release, scheduling and change controls
- Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning
- Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R
- Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports
- ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation
- Reconciliation, data quality and SLA monitoring
- Performance at scale: tables of 1 billion+ rows and multi-year history
Requirements
Experience
- Lead: 12+ years, including 6+ in banking.
- Senior: 8+ years, including 4+ in banking
Banking and domain
- Production delivery of data pipelines inside banks
- Customer, account, transaction, payments and AML data
- Working within bank release, scheduling and change controls
- Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning
- Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R
- Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports
- ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation
- Reconciliation, data quality and SLA monitoring
- Performance at scale: tables of 1 billion+ rows and multi-year history
Core data engineering
- Production delivery of data pipelines inside banks
- Customer, account, transaction, payments and AML data
- Working within bank release, scheduling and change controls
- Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning
- Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R
- Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports
- ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation
- Reconciliation, data quality and SLA monitoring
- Performance at scale: tables of 1 billion+ rows and multi-year history
Integration and platforms
- Kafka, Spark Streaming, Informatica (PowerCenter, IDMC, IDL) and Talend; REST API development
- Denodo, Snowflake and NoSQL databases
- Airflow, Control-M or Autosys; Git, CI/CD, GitOps and Kubernetes
Delivery and communication (Lead)
- Framework design, code standards, code reviews and estimation
- Guiding a team of engineers and working with architects and analysts
Good-to-have Skills
- Databricks: Delta Lake, Unity Catalog and Workflows
- Data services on Azure, Google Cloud, Huawei Cloud or Alibaba Cloud
- Data modelling for Qlik Sense or Power BI
Certifications (preferred)
Teradata Vantage; Cloudera Data Engineer; Databricks Data Engineer Associate or Professional