An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Jobtailor in Bengaluru is seeking an experienced data engineer to design, build, and scale batch and streaming pipelines using PySpark, Azure Databricks, and Microsoft Fabric. You will implement the Medallion Architecture across Bronze, Silver, and Gold layers to support enterprise analytics and data products.
You will optimize Spark workloads, ensure data quality, governance, and lineage, and collaborate with architects and product teams in an Agile environment.
• Design, develop, and maintain scalable batch and streaming data pipelines using PySpark, Azure Databricks, and Microsoft Fabric
• Architect and implement Medallion Architecture across Bronze, Silver, and Gold layers for enterprise data processing and analytics
• Build robust ETL/ELT solutions for data ingestion, transformation, validation, reconciliation, and delivery across multiple source systems
• Develop and optimize PySpark and Spark SQL workloads for high-volume structured, semi-structured, and unstructured data
• Design and maintain Lakehouse and data lake solutions using Azure Data Lake Storage Gen2, Delta Lake, Microsoft Fabric OneLake, Fabric Lakehouse, and Warehouse
• Implement integration solutions using Azure Data Factory, Fabric Data Factory, data pipelines, notebooks, and Dataflows Gen2
• Design secure and governed data-sharing solutions across workspaces, domains, business units, and approved external consumers
• Implement reusable data products and cross-workspace sharing patterns using OneLake, OneLake shortcuts, Lakehouse, Warehouse, and semantic models
• Contribute to Microsoft Fabric capacity planning, workspace-to-capacity assignment, workload monitoring, utilization analysis, and performance optimization
• Monitor Fabric workloads using available capacity and workload metrics, identify resource contention, and recommend workload or scheduling improvements
• Design domain-aligned Fabric workspace structures with appropriate separation for development, testing, production, security, and ownership boundaries
• Implement data quality controls, monitoring, observability, lineage, error handling, reconciliation, and auditability across data pipelines
• Apply security best practices using managed identities, role-based access control, workspace roles, row-level or object-level controls where applicable, and secure secrets management
• Integrate data engineering solutions with Git-based source control and CI/CD pipelines for automated testing and deployment
• Optimize performance, scalability, reliability, and cost across Azure Databricks and Microsoft Fabric workloads
• Collaborate with Data Architects, Product Owners, Business Analysts, Data Scientists, QA engineers, governance teams, and platform teams in an Agile/Scrum environment
• Provide technical leadership, conduct design and code reviews, establish engineering standards, and mentor data engineers
Demonstrates expertise in designing and implementing scalable data pipelines and architectures using PySpark, Azure Databricks, and Microsoft Fabric. Proficient in ETL/ELT processes, data quality management, and performance optimization across cloud data platforms.