Design, develop, and deploy enterprise-scale data engineering solutions using Databricks and Azure. The role focuses on building robust and scalable data pipelines, metadata-driven ingestion frameworks, data quality solutions, and Lakehouse architectures. The ideal candidate will collaborate with global cross-functional teams, provide technical leadership, mentor junior engineers, and drive high-quality delivery across the full project lifecycle.
Roles and Responsibilities
- Data Engineering & Architecture: Design and develop scalable data solutions aligned with technical architecture, integration standards, and established engineering practices.
- Pipeline Development: Build and optimize enterprise-grade data pipelines using Databricks, PySpark, Delta Lake, Delta Live Tables, Auto Loader, Databricks Workflows, and/or Apache Airflow.
- Metadata & Data Quality: Develop metadata-driven ingestion frameworks and robust data quality (DQ) frameworks using PySpark to ensure reliable and governed data processing.
- Lakehouse & Data Platforms: Design and implement modern Lakehouse architectures using Apache Spark, Delta Lake, Azure Data Lake Storage, and related big data technologies.
- Technical Leadership: Lead technical discussions, clarify requirements, resolve ambiguities, conduct design and code reviews, and provide technical guidance and mentorship to junior team members.
- Performance & Optimization: Perform performance tuning and optimization across Databricks and Apache Spark workloads, pipelines, queries, and data processing jobs.
- Security & Governance: Implement Databricks Unity Catalog and fine-grained access controls to support enterprise data governance and security requirements.
- Client & Cross-Functional Collaboration: Work closely with business analysts, functional teams, onsite clients, architects, and global delivery teams to translate requirements into effective technical solutions.
- Automation & Delivery: Develop reusable templates, frameworks, and scripts to automate development and operational activities. Support project estimation, planning, execution, tracking, and continuous improvement.
Required Skills
- Databricks, Azure, Python, SQL, PySpark, Apache Spark, Delta Lake, Lakehouse Architecture, Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Data Factory, Azure Synapse, Azure Key Vault, Cosmos DB, Unity Catalog, Delta Live Tables, Auto Loader, Databricks Workflows, and Apache Airflow.
- 5–7 years of experience in Data Engineering with significant hands‑on expertise in Databricks on Azure.
- Strong experience building metadata-driven ingestion and data quality frameworks using PySpark.
- Strong understanding of Lakehouse architecture, data lakes, data warehouses, data marts, 3NF, dimensional modeling, and modern data platforms.
- Hands‑on experience with Databricks/Spark performance tuning, data pipeline optimization, and scalable data processing.
- Experience implementing Unity Catalog and fine-grained data access controls.
- Strong problem-solving, analytical, communication, stakeholder management, and team collaboration skills.
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status.