A leading company is seeking a Data Engineer to design and implement data solutions on Databricks. The role involves developing ETL pipelines, managing data quality, and collaborating with stakeholders to meet analytics needs. Ideal candidates will have strong skills in Python, Pyspark, and AWS services.
Qualifications
Experience with data modeling and ETL pipelines using cloud-based services.
Experience with Collibira for data quality and governance is a plus.
Responsibilities
Design data solutions on Databricks including delta lake and data warehouse.
Develop and manage data engineering and ETL processes.
Interact with stakeholders for data landscape understanding.
Skills
Databricks
Pyspark
Python
SQL
AWS Services
Collibira DQ
Unity catalog
Job description
Mandatory Skills
Databricks – Pyspark,Python Job,SQL
Unity catalog
Collibira DQ
AWS Services ex. Glue,s3,redshift,lambda etc
Job Description:
Design of data solutions on Databricks including delta lake, data warehouse, data marts and other data solutions to support the analytics needs of the organization.
Apply best practices during design in data modeling (logical, physical) and ETL pipelines (streaming and batch) using cloud-based services especially Python & Pyspark
Design, develop and manage the pipelining (collection, storage, access), data engineering (data quality, ETL, Data Modelling) and understanding (documentation, exploration) of the data.
Interact with stakeholders regarding data landscape understanding, conducting discovery exercises, developing proof of concepts, and demonstrating it to stakeholders.
Experience to work on Collibira for DQ and data governance is plus
Knowledge on dbt to model and build out layers in dbx is plus
AWS experience on s3,redshift,glue etc is also required