An application made for this job — a tailored resume and cover letter that speak straight to the posting.
EXL is seeking a Senior Data Engineer to build and operate data pipelines feeding the Entity Hub on Microsoft Fabric. You will land six sources into the Fabric Bronze/raw layer, implement standardization and transformation logic, and ensure data quality with monitoring and lineage.
The role emphasizes PySpark, SQL, Delta Lake and CDC patterns, with strong focus on ingestion pipelines, data quality, and performance tuning within a Fabric-based lakehouse architecture.
Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programme rests on.
Data Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts
Batch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats
Validation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation
Bronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes
Pipeline monitoring, lineage and metadata capture, access controls, technical documentation
Complementary with the Entity Resolution engineering workstream — both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.