Job Type: Contract
Work Mode: Remote
Mode: 100% remote
Time Shift: EST Time shift
Pref: Immediate Joiner
We are looking for an experienced Databricks + Azure Developer to design, build, and optimize scalable data engineering solutions on Azure. The ideal candidate will have strong hands-on experience with Azure Databricks, PySpark, Spark SQL, Delta Lake, Azure Data Factory, ADLS Gen2, and modern data-pipeline development. The role involves building high-performance ETL/ELT pipelines, implementing Lakehouse architecture, and ensuring reliable, secure, and scalable data operations.
Key Responsibilities
- Design, develop, and maintain data pipelines using Azure Databricks and PySpark.
- Build and orchestrate ETL/ELT workflows using Azure Data Factory (ADF).
- Develop ingestion and transformation pipelines for structured and semi-structured data using ADLS Gen2, Delta Lake, and Spark SQL.
- Implement Medallion architecture (Bronze/Silver/Gold layers) and incremental loads, CDC, MERGE, and SCD Type 1/2.
- Optimize Spark jobs through partitioning, caching, file-size optimization, and performance-tuning techniques.
- Configure and maintain Databricks Workflows/Jobs, clusters, and workspace environments.
- Implement data quality checks, validation, error handling, monitoring, and reconciliation processes.
- Work with Unity Catalog for governance, lineage, RBAC, and secure data access.
- Integrate Databricks with Azure services (ADLS, Azure SQL, APIs, JDBC connections).
- Implement CI/CD pipelines using Azure DevOps and Git; automate infrastructure using Terraform or ARM/Bicep.
- Troubleshoot production pipelines, perform root-cause analysis, and ensure SLA compliance.
- Collaborate with architects, analysts, BI teams, and business stakeholders in an Agile environment.
Required Qualifications
- 9–10+ years of experience in Data Engineering or Azure cloud data development.
- Strong hands-on experience with Azure Databricks, PySpark, Spark SQL, and Delta Lake.
- Experience building scalable ETL/ELT pipelines using Azure Data Factory (ADF).
- Strong understanding of ADLS Gen2, Lakehouse architecture, and distributed data processing.
- Experience with CI/CD (Azure DevOps), Git, YAML pipelines, and Infrastructure-as-Code (Terraform/ARM/Bicep).
- Solid SQL skills and experience with performance tuning.
- Experience working in Agile/Scrum environments.
Preferred Qualifications
- Experience with Databricks platform administration (clusters, jobs, workspace configuration).
- Knowledge of Unity Catalog, data governance, lineage, and RBAC.
- Experience integrating Databricks with APIs, JDBC sources, and external systems.
- Familiarity with monitoring tools (Datadog, Grafana) and cost optimization.