Data Engineer Fabric

EXL

Gurugram District

On-site

INR 900,000 - 1,300,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

EXL is seeking a Senior Data Engineer to build and operate data pipelines feeding the Entity Hub on Microsoft Fabric. You will land six sources into the Fabric Bronze/raw layer, implement standardization and transformation logic, and ensure data quality with monitoring and lineage.

The role emphasizes PySpark, SQL, Delta Lake and CDC patterns, with strong focus on ingestion pipelines, data quality, and performance tuning within a Fabric-based lakehouse architecture.

Qualifications

  • 4+ years hands-on data engineering with PySpark and SQL.
  • Production experience building ingestion pipelines from multiple heterogeneous sources.
  • Working knowledge of Delta Lake and medallion/lakehouse architecture.
  • Experience implementing incremental loads and CDC-style processing.
  • Experience implementing data quality checks and troubleshooting pipeline failures.

Responsibilities

  • Ingestion development — build and maintain pipelines to land six in-scope sources into Fabric Bronze/raw layer.
  • Mirroring & CDC — implement Fabric Mirroring and change-data-capture patterns; watermark/incremental load logic.
  • Raw layer management — maintain one Delta table per source, append-only with provenance.
  • Standardization & transformation — implement name normalization, address parsing in Spark notebooks; support identifier-spine construction.
  • Data quality — implement checks, validation rules, alerts, and reconciliation against source.
  • Pipeline operations — schedule, monitor, troubleshoot runs; maintain run docs.
  • Performance tuning — optimize Spark jobs, Delta file sizes, and partitioning for Fabric capacity.
  • Documentation — produce and maintain source-to-target mappings and lineage records.

Skills

PySpark
SQL
Delta Lake
CDC patterns
Data quality checks
Troubleshooting pipelines
Microsoft Fabric
Fabric ingestion pipelines

Tools

Microsoft Fabric
Spark notebooks
Delta Lake
Great Expectations

Job description

Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programme rests on.

Key Responsibilities
  • Ingestion development — build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.
  • Mirroring & CDC — implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.
  • Raw layer management — maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.
  • Standardization & transformation — implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.
  • Data quality — implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.
  • Pipeline operations — schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.
  • Performance tuning — optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.
  • Documentation — produce and maintain source-to-target mappings, transformation logic documentation and lineage records.
Required Skills & Experience
Skill Area
Specific Requirements
Core Engineering
Microsoft Fabric

Data Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts

Data Integration

Batch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats

Data Quality

Validation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation

Bronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes

Ops & Governance

Pipeline monitoring, lineage and metadata capture, access controls, technical documentation

Must-Have Qualifications
  • 4+ years hands-on data engineering with strong PySpark and SQL
  • Production experience building ingestion pipelines from multiple heterogeneous sources
  • Working knowledge of Delta Lake and medallion/lakehouse architecture
  • Experience implementing incremental loads and CDC-style processing
  • Experience implementing data quality checks and troubleshooting pipeline failures
Nice-to-Have
  • Microsoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)
  • Exposure to entity/master data standardization (name and address parsing)
  • Familiarity with libraries such as Great Expectations for data quality
  • Experience optimising for Fabric capacity/CU consumption
  • Operational ingestion pipelines for all agreed sources
  • Bronze/raw layer with one Delta table per source and CDC retained
  • Standardization and parsing transformation logic
  • Data quality checks, monitoring and exception handling
  • Source-to-target mapping and run documentation
Dual Role / Complementary Skills

Complementary with the Entity Resolution engineering workstream — both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Fabric Senior Data Engineer
Fabric Senior Data Engineer

EXL • Pune District

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Protiviti India • Hyderabad, Pune District, Coimbatore District

Hybrid
INR 900,000 - 1,500,000
Data Engineer with Fabric
Data Engineer with Fabric

Airo Digital Labs, LLC • India

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Y&L Consulting • Indore District, Hyderabad, Pune District

Hybrid
INR 1,400,000 - 1,800,000
Data Engineer
Data Engineer

TalentBridge • India

On-site
INR 900,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Version 1 • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Lead Data Engineer - Fabric
Lead Data Engineer - Fabric

iLink Digital • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Data & AI Engineer - Microsoft Fabric
Senior Data & AI Engineer - Microsoft Fabric

Exponentia.ai • Mumbai

On-site
INR 1,800,000 - 3,500,000
Senior Data Engineer Azure and Microsoft Fabric
Senior Data Engineer Azure and Microsoft Fabric

Sonata Software • Hyderabad, Chennai District, Bengaluru

Hybrid
INR 2,600,000 - 4,200,000
Microsoft Fabric Data Engineer
Microsoft Fabric Data Engineer

Ecolab • Bengaluru

On-site
INR 1,100,000 - 1,700,000