An innovative firm is seeking a skilled Data Engineer to develop and optimize ETL/ELT pipelines using PySpark and SQL. In this role, you will work with both structured and unstructured data to build scalable data solutions, ensuring data quality and performance optimization. Collaborating closely with Data Scientists and Analysts, you will design data models and implement robust data workflows on cloud platforms. This position offers a fantastic opportunity to contribute to cutting-edge data projects and enhance your skills in a dynamic and supportive environment.
Qualifications
Strong experience in Python and PySpark for data engineering tasks.
Deep understanding of SQL and ETL/ELT development using Spark.
Responsibilities
Develop and maintain ETL/ELT pipelines using PySpark and SQL.
Collaborate with Data Scientists and Analysts to integrate data workflows.
Skills
Python
PySpark
SQL
ETL/ELT Development
Cloud Data Services
Data Warehousing
Orchestration Tools
Version Control (Git)
Tools
AWS Glue
Databricks
Azure Synapse
GCP BigQuery
Airflow
Apache Oozie
Snowflake
Redshift
Job description
Job Responsibilities:
Develop, optimize, and maintain ETL/ELT pipelines using PySpark and SQL.
Work with structured and unstructured data to build scalable data solutions.
Write efficient and scalable PySpark scripts for data transformation and processing.
Optimize SQL queries, stored procedures, and indexing strategies to enhance performance.
Design and implement data models, schemas, and partitioning strategies for large-scale datasets.
Collaborate with Data Scientists, Analysts, and other Engineers to integrate data workflows.
Ensure data quality, validation, and consistency in data pipelines.
Implement error handling, logging, and monitoring for data pipelines.
Work with cloud platforms (AWS, Azure, or GCP) for data processing and storage.
Optimize data pipelines for cost efficiency and performance.
Technical Skills Required:
Strong experience in Python for data engineering tasks.
Proficiency in PySpark for large-scale data processing.
Deep understanding of SQL (Joins, Window Functions, CTEs, Query Optimization).
Experience in ETL/ELT development using Spark and SQL.
Experience with cloud data services (AWS Glue, Databricks, Azure Synapse, GCP BigQuery).
Familiarity with orchestration tools (Airflow, Apache Oozie).
Experience with data warehousing (Snowflake, Redshift, BigQuery).
Understanding of performance tuning in PySpark and SQL.
Familiarity with version control (Git) and CI/CD pipelines.