Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)

HR POD - Hiring Talent Globally

Lahore

On-site

PKR 300,000 - 600,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

HR POD - Hiring Talent Globally seeks a data engineer to own end-to-end ingestion pipelines using Databricks, Delta Lake and Medallion Architecture. You will work with Spark SQL, PySpark, and production-grade streaming data from Azure Event Hubs, Kafka, or Kinesis.

Responsibilities include handling schema drift, data quality, and coordinating with a US-based team to resolve incidents and ensure data correctness across Bronze-Silver-Gold layers.

Qualifications

  • 4+ years of experience in data engineering with production Databricks experience.
  • Strong expertise in Spark SQL, PySpark, Delta Lake, Medallion Architecture and Delta Live Tables (DLT).
  • Hands-on with streaming ingestion using Azure Event Hubs, Kafka or Kinesis.
  • Experience debugging data discrepancies across source-to-warehouse pipelines and handling schema drift.

Responsibilities

  • Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
  • Work with voice-AI observability and telephony data platforms.
  • Handle multiple pagination schemes and manage rate limits with time-windowing.
  • Implement robust schema-drift handling and alert when fields are renamed or removed.
  • Own streaming ingestion from Azure Event Hubs into Databricks with Structured Streaming/Auto Loader.
  • Manage checkpoints, offsets, watermarking and at-least-once deduplication.
  • Develop Delta Lake pipelines using Medallion Architecture (Bronze-Silver-Gold).
  • Build entity-resolution pipelines for dirty, free-text data.

Skills

SQL
Python
Data pipelines
Data quality

Tools

Databricks
Spark SQL
PySpark
Delta Lake
Delta Live Tables
Azure Event Hubs
Kafka
Kinesis
Holistics AML/AQL
dbt Metrics

Job description

Requirements
  • 4+ years of experience in data engineering, with substantial production experience in Databricks.
  • Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture, and Delta Live Tables (DLT).
  • Hands-on experience with Structured Streaming or equivalent production-grade streaming ingestion using Azure Event Hubs, Kafka, or Kinesis.
  • Strong understanding of checkpoint recovery, watermarking, and deduplication strategies.
  • Demonstrated experience debugging source-to-warehouse data discrepancies.
  • Ability to walk through a real-world incident involving mismatched record counts and explain how the root cause was identified and resolved.
  • Proven experience integrating third-party REST APIs in production.
  • Experience handling pagination edge cases, rate and row limits, retries, and schema drift.
  • Experience with entity resolution or data matching involving messy, real-world text data.
  • Experience with a metrics or semantic layer such as Holistics AML/AQL, dbt Metrics, or LookML.
  • Working understanding of why non-additive measures cannot be reliably calculated from pre-aggregated rollups.
  • Strong SQL and Python skills, with the ability to own data pipelines end-to-end with minimal oversight.
  • Strong written and spoken English, with the ability to collaborate effectively with a US-based team asynchronously.
  • Healthcare data experience, including referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets.
  • Familiarity with HIPAA handling expectations.
  • Experience with voice-agent, call-center, or telephony/conversation data.
  • Familiarity with call transcripts and containment or outcome metrics.
  • Hands-on experience with Holistics, specifically AML/AQL modeling.
  • Experience with the broader Azure ecosystem beyond Event Hubs, including ADLS, ADF, and Key Vault.
Responsibilities
  • Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
  • Work primarily with voice-AI observability and telephony platforms.
  • Handle different pagination schemes, including offset/limit and page/cursor models.
  • Manage row and rate limits through time-windowing and adaptive bisection.
  • Implement robust schema-drift handling through contract and column-presence checks.
  • Ensure alerts are triggered when fields are renamed, moved, or removed rather than silently propagating null values.
  • Own streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and/or Auto Loader.
  • Manage checkpoints and offsets, watermarking, and at-least-once deduplication.
  • Perform source-parity reconciliation across ingested and production data.
  • Investigate row counts, dropped or duplicated events, late-arriving data, and schema mismatches.
  • Identify and resolve the root cause when ingested data does not match production sources.
  • Develop and maintain Delta Lake pipelines using a Medallion Architecture (Bronze Silver Gold).
  • Use Spark SQL and PySpark to build and maintain production data pipelines.
  • Implement idempotent MERGE upserts.
  • Work with Delta Live Tables and materialized-view constraints, including CREATE OR REFRESH and LIVE references.
  • Understand and manage differences between DLT and job execution contexts.
  • Build entity-resolution pipelines for dirty, free-text data.
  • Normalize practice, provider, and payer names using regex, canonical dictionaries, fuzzy matching, confidence-scored crosswalks, and override tables.
  • Maintain the semantic and metrics layer with rigorous metric definitions.
  • Define and maintain accurate denominators, data grain, and cohort boundaries.
  • Ensure the correct handling of non-additive aggregates, including medians and percentiles that cannot be reliably supported through aggregate-aware pre-aggregation.
  • Ensure every metric remains accurate and reproducible.
  • Instrument data quality across the entire pipeline.
  • Monitor data freshness, source parity, data contracts, and other critical quality checks.
  • Build alerting mechanisms that identify data issues before they reach dashboards.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Databricks Engineer
Senior Databricks Engineer

Digifloat • Islamabad

On-site
PKR 2,600,000 - 3,400,000
Senior Databricks Data Engineer: Streaming & Delta Lake
Senior Databricks Data Engineer: Streaming & Delta Lake

HR POD - Hiring Talent Globally • Lahore

On-site
PKR 300,000 - 600,000
Data Engineer
Data Engineer

Zorba Consulting • Hyderabad City Taluka

On-site
INR 1,200,000 - 2,400,000
Principal Software Engineer (Data Platform)
Principal Software Engineer (Data Platform)

NorthBay Solutions • Islamabad

Hybrid
PKR 1,200,000 - 2,100,000
Data Engineer
Data Engineer

Acme One • Lahore

On-site
PKR 1,500,000 - 3,000,000
Principal Software Engineer (Data Platform)
Principal Software Engineer (Data Platform)

NorthBay Solutions • Pakistan

Hybrid
PKR 300,000 - 420,000
Data Management & Analytics Specialist
Data Management & Analytics Specialist

NorthBay Solutions • Islamabad

Hybrid
PKR 1,800,000 - 2,800,000
Hybrid work model
Data Management & Analytics Specialist
Data Management & Analytics Specialist

NorthBay Solutions LLC • Lahore

Hybrid
PKR 2,400,000 - 4,000,000
Principal Software Engineer (Data Platform)
Principal Software Engineer (Data Platform)

NorthBay Solutions • Lahore

On-site
PKR 2,000,000 - 4,200,000
Data Management & Analytics Specialist
Data Management & Analytics Specialist

NorthBay Solutions LLC • Pakistan

Hybrid
PKR 250,000 - 500,000