Key Responsibilities
- Design and deliver endtoend data solutions using Microsoft Fabric including OneLake Lakehouse Data Pipelines Fabric Notebooks and Power BI semantic models
- Build scalable batch incremental and nearrealtime ingestion pipelines from relational filebased API and eventdriven sources
- Develop reusable transformations using Python PySpark and SQL with modular patterns for data cleansing enrichment reconciliation and businessrule processing
- Implement medallionstyle data architecture using Bronze Silver and Gold layers Delta tables and analyticsready curated datasets
- Design dimensional models fact and dimension tables semantic models and governed business metrics for reporting and analytics
- Configure data quality checks schema validation deduplication replay backfill and errorhandling patterns to improve pipeline reliability
- Implement security and access controls using Microsoft Entra ID RBAC service principals workspace roles leastprivilege principles and appropriate sensitivedata controls
- Enable metadata management lineage classification and governance in collaboration with enterprise data governance and security teams
- Create CICD and sourcecontrol practices for Fabric artifacts and related data engineering assets using Git or GitHub Actions
- Monitor pipeline health troubleshoot production failures perform root cause analysis and tune workloads for performance capacity and cost efficiency
- Partner with architects analysts product owners and business stakeholders to translate requirements into secure supportable and welldocumented data products
- Contribute to coding standards design reviews peer reviews reusable patterns technical documentation and knowledge transfer
Required Qualifications
- 3 years of experience in data engineering data platforms analytics engineering or a closely related role
- Demonstrated handson experience delivering production solutions with Microsoft Fabric
- Strong practical knowledge of OneLake Lakehouse Data Pipelines Fabric Notebooks and Power BI semantic models
- Advanced development skills in Python PySpark and SQL including performance tuning and troubleshooting
- Experience designing ETLELT pipelines Delta Lake or equivalent lakehouse patterns and dimensional data models
- Experience with data quality observability logging ing and production support for enterprise data pipelines
- Experience implementing identity access control secrets management and secure servicetoservice connectivity in Azure environments
- Working knowledge of Gitbased version control and automated deployment practices for data platform assets
- Strong communication skills and the ability to work effectively with technical and business stakeholders
Preferred Qualifications
- Microsoft Certified Fabric Data Engineer Associate or an equivalent Microsoft data certification
- Experience with ADLS Gen2 or Azure streaming services
- Experience with Delta Lake Kafka Event Hubs Fivetran or comparable ingestion and transformation technologies
- Experience implementing data cataloging lineage classification privacy controls and regulatory datagovernance requirements
- Experience with Terraform and broader DevSecOps practices
- Exposure to crosscloud data platforms such as Snowflake and AWS is beneficial
- Experience coaching engineers conducting design reviews and defining reusable engineering standards
Technical Competency Profile
- AreaExpected capability
- Core platformMicrosoft Fabric OneLake Lakehouse Data Pipelines Fabric Notebooks
- EngineeringPython PySpark SQL ETLELT batch and streaming ingestion
- AnalyticsPower BI Semantic models dimensional modeling governed business metrics
- Azure ecosystemADLS Gen2 Event Hubs
- Security governanceMicrosoft Entra ID RBAC service principals workspace roles lineage classification
- DevOps operationsGit GitHub Actions CICD monitoring ing RCA performance tuning
Success Measures
- Reliable and timely delivery of secure analyticsready datasets and data products
- Improved pipeline stability observability recoverability and adherence to agreed service expectations
- Reusable engineering patterns effective environment promotion and reduced manual deployment effort
- Measurable improvements in workload performance data quality governance and platform cost efficiency
- Clear documentation effective stakeholder communication and successful knowledge transfer