Location: Krakow, Poland (Hybrid Model – On-site Presence Required)
Number of Positions: 3
Experience: 6–10 years (preferably 8–10 years)
Job Overview:
- We are seeking experienced Data Engineers with strong expertise in PySpark and Python to join our growing data engineering team in Kraków.
- The successful candidates will be responsible for designing, developing, and optimizing scalable data pipelines and data processing solutions for large-scale enterprise environments.
- Experience with Azure Data Factory (ADF) is required, while hands-on experience with Azure Databricks (ADB) will be considered a significant advantage.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using PySpark and Python.
- Build and optimize ETL/ELT workflows for processing large volumes of data.
- Write efficient, reusable, and high-performance code with a strong focus on optimization.
- Process and manage structured and unstructured data from multiple sources.
- Implement best practices related to data engineering, performance tuning, and code quality.
- Develop, orchestrate, and monitor data workflows using Azure Data Factory (ADF).
- Collaborate with Data Scientists, Data Analysts, Solution Architects, and other stakeholders.
- Ensure data quality, reliability, and integrity across all data platforms.
- Troubleshoot and resolve performance bottlenecks in data processing pipelines.
- Contribute to data architecture discussions and continuous improvement initiatives.
Required Skills & Experience
- 6–10 years of overall IT experience.
- 4–5 years of hands-on experience in PySpark and Python development.
- Strong understanding of data engineering principles and modern data pipeline architectures.
- PySpark (Spark SQL, DataFrames, performance tuning)
- Python scripting and application development
- Azure Data Factory (ADF) orchestration and pipeline development
- Strong knowledge of data optimization techniques, including partitioning, caching, joins, and query optimization.
- Experience working with large-scale distributed data processing systems.
- Strong analytical and problem-solving skills.
- Good understanding of SQL, database concepts, and data modeling.
Nice-to-Have Skills
- Hands-on experience with Azure Databricks (ADB).
- Exposure to cloud technologies, preferably Microsoft Azure.
- Experience with CI/CD pipelines and DevOps practices for data engineering.
- Knowledge of Data Lake and Lakehouse architectures.
- Familiarity with modern data governance and monitoring practices.