A data solutions company is seeking an experienced data engineer to design data solutions on Databricks. In this role, you will apply best practices in data modeling and ETL pipelines using Python and AWS services. Key responsibilities include managing data pipelines and collaborating with stakeholders on data governance. Knowledge of Collibra and dbt is a plus.
Qualifications
Experience designing data solutions using Databricks.
Proficient in Python and Pyspark for ETL.
Knowledge of AWS services like S3, Redshift, and Glue.
Responsibilities
Design data solutions to meet analytics needs.
Develop and manage data pipelines and engineering processes.
Engage with stakeholders for data governance and modeling.
Skills
Data modeling
ETL pipelines
Python
Pyspark
AWS services
Tools
Databricks
Collibra
dbt
Job description
Responsibilities
Design of data solutions on Databricks including delta lake, data warehouse, data marts and other data solutions to support the analytics needs of the organization.
Apply best practices during design in data modeling (logical, physical) and ETL pipelines (streaming and batch) using cloud-based services especially Python & Pyspark
Design, develop and manage the pipelining (collection, storage, access), data engineering (data quality, ETL, Data Modelling) and understanding (documentation, exploration) of the data.
Interact with stakeholders regarding data landscape understanding, conducting discovery exercises, developing proof of concepts, and demonstrating it to stakeholders.
Experience to work on Collibira for DQ and data governance is plus
Knowledge on dbt to model and build out layers in dbx is plus
AWS experience on s3,redshift,glue etc is also required