We are looking for an experienced Azure Databricks Engineer to design, develop, and optimize scalable data engineering solutions using Azure Databricks, PySpark, Azure Data Factory, ADLS Gen2, SQL, and Delta Lake.
The ideal candidate should have strong hands-on experience building enterprise-grade ETL/ELT pipelines, implementing Lakehouse architecture, optimizing Spark workloads, and working with large-scale structured and semi-structured datasets.
Key Responsibilities
- Design and develop scalable data pipelines using Azure Databricks and PySpark.
- Build and maintain ETL/ELT pipelines using Azure Data Factory (ADF).
- Develop data ingestion and transformation workflows for structured and semi-structured data.
- Work extensively with ADLS Gen2, Delta Lake, Azure SQL, and Azure Databricks.
- Implement Bronze, Silver, and Gold / Medallion architecture.
- Develop complex transformations using PySpark and Spark SQL.
- Implement incremental loads, CDC, MERGE operations, and SCD Type 1/Type 2.
- Optimize Spark jobs through partitioning, caching, file-size optimization, and other performance-tuning techniques.
- Implement and maintain Databricks Workflows/Jobs for pipeline orchestration.
- Implement data quality, validation, error handling, monitoring, and reconciliation processes.
- Work with Unity Catalog for data governance, access control, lineage, and security.
- Implement CI/CD pipelines using Azure DevOps and Git.
- Troubleshoot production data pipelines and perform root-cause analysis.
- Collaborate with data architects, analysts, BI teams, and business stakeholders.
- Follow coding, documentation, security, and data engineering best practices.
Required Technical Skills
- 10+ years of experience in Data Engineering.
- Strong hands-on experience with Azure Databricks.
- Strong expertise in PySpark / Apache Spark.
- Advanced SQL skills.
- Hands-on experience with Azure Data Factory (ADF).
- Strong experience with Azure Data Lake Storage Gen2 (ADLS Gen2).
- Strong knowledge of Delta Lake.
- Experience with Medallion/Lakehouse architecture.
- Experience with ETL/ELT development and data pipeline design.
- Experience with Databricks Jobs/Workflows.
- Knowledge of Unity Catalog and data governance.
- Experience with Git and Azure DevOps / CI-CD.
- Strong understanding of data modeling and data warehousing concepts.
Good to Have
- Experience with Delta Live Tables (DLT).
- Experience with Auto Loader.
- Knowledge of Structured Streaming.
- Experience with Azure Synapse Analytics.
- Experience with Azure Key Vault and Managed Identity.
- Experience with Kafka / Event Hubs.
- Experience with Terraform or Infrastructure as Code.
- Experience with Power BI or other BI platforms.
- Databricks or Microsoft Azure certifications.
Preferred Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Strong problem-solving and analytical skills.
- Good communication and collaboration skills.
- Ability to work independently in a fast-paced environment.
Core Technology Stack
Databricks: PySpark, Spark SQL, Delta Lake, Unity Catalog, Workflows, DLT, Auto Loader
Programming: Python, SQL
Architecture: Lakehouse, Medallion Architecture, ETL/ELT, Data Warehousing