Bringing data closer to people. ScatterPie Analytics is a company focused on modernizing data analytics through the implementation of technologies that companies need today to succeed in the market. We bring together the best of open source and blend with the best of enterprise tools to build intelligence that is powerful, cost-effective and flexible.
Job Description
Key Responsibilities:
- Design, develop, and optimize data pipelines using Databricks, PySpark, Delta Lake, and related Azure components.
- Work across pipeline development, data modeling, and production code optimization, ensuring scalability and performance.
- Build new pipelines from scratch as well as enhance and maintain existing ones.
- Apply strong understanding of ETL design principles, data modeling concepts, and data governance standards.
- Collaborate within a scrum team, taking ownership of assigned stories while independently delivering high-quality, production-ready code.
- Demonstrate proficiency across PySpark, SQL, Delta Lake, Unity Catalog, and Databricks Workflows, with solid understanding of logic and data flow.
- Work in Azure environments, leveraging tools like ADLS, ADF, and Synapse (as applicable).
- Contribute inputs to architecture and design discussions where appropriate, while primarily focusing on hands-on development.
- Troubleshoot data issues, perform root-cause analysis, and optimize performance for existing pipelines.
Requirements
Desired Skills and Experience:
- 2-3 years of experience in data engineering, preferably with Azure Databricks.
- Strong technical expertise in PySpark, SQL, and Delta Lake.
- Familiarity with Unity Catalog, data governance, and DevOps practices (Git, CI/CD).
- Ability to work independently as well as collaboratively within an agile delivery team.
- Excellent problem-solving, debugging, and communication skills.