Senior Data Engineer

Quantiphi Analytics Solutions Private Limited

Mumbai

On-site

INR 2,000,000 - 4,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Quantiphi Analytics Solutions Private Limited is hiring a Data Engineer in Mumbai with 4–8 years of experience. The role focuses on building ingestion pipelines with Flink and PySpark, implementing HL7v2, CCDA, CSV, and FHIR parsers, and developing the dbt transformation layer across CDM models.

You will work on FHIR serialization, FHIR-Repository integrations, and orchestration with Cloud Composer within a hybrid, collaborative team environment.

Qualifications

  • Strong Python and SQL proficiency.
  • 4+ years of hands-on data engineering experience.
  • Experience with PySpark and PyFlink preferred.

Responsibilities

  • Build and maintain ingestion pipelines with Flink streaming and PySpark batch jobs.
  • Implement source-format parsers (HL7v2, CCDA, CSV, FHIR) as Python classes.
  • Write tests covering input scenarios and DLQ routing.
  • Maintain synchronous MDM call patterns and DLQ routing for failures.
  • Develop dbt transformations across staging, intermediate, and CDM models.
  • Create and maintain dbt macros and a FHIR serialization layer.
  • Collaborate on Cloud Composer DAGs and CICD pipelines.

Skills

Python
SQL
PySpark
PyFlink
Java/Kotlin exposure
GCP
DBT
Kafka
Flink
HL7v2/FHIR knowledge
Git CI/CD
Debugging/problem solving

Education

Bachelor's or Master's in CS/Engineering

Tools

Dataproc
Cloud Composer
GKE/Cloud Run
dbt
Kafka
Flink/Spark
Git

Job description

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth. If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Data Engineer Exp Range : 4 - 8 Years Location : Mumbai , Bangalore, Trivandrum

Role review :

The Data Engineer is the implementation backbone of the platform.

Key Responsibilities
  • Build and maintain the ingestion pipelines — Apache Flink streaming jobs on Dataproc for HL7v2 and FHIR feeds, PySpark batch jobs for CCDA and CSV bulk loads.
  • Implement against the CDM_ingest_mapping shared Python library that defines SourceSpec instances per source × format combination.
  • Implement source-format parsers (HL7v2, CCDA, CSV, FHIR) as Python classes per the parser component spec.
  • Write fixture-driven tests covering both well-formed inputs and DLQ-routing scenarios.
  • Implement and maintain the synchronous Informatica MDM call pattern in the ingestion path — batched calls, timeout handling, circuit breaker behavior, DLQ routing for MDM failures.
  • Implement the asynchronous MDM event consumer that applies ECI changes to CDM.
  • Build the dbt transformation layer end to end — staging models per source, intermediate models that union and resolve REFs, CDM target models (DIM/FACT/BRIDGE/REF) that apply SCD2 via shared macros, and data product models.
  • Write the dbt YAML schemas, tests, and documentation that accompany every model.
  • Implement and maintain the shared dbt macro library — hash_key, scd2_merge, attribute_hash, restate_merge, audit_columns. Macros are the most-reused code; their correctness is non-negotiable and they require golden tests.
  • Build the FHIR serialization layer — flat FHIR Iceberg tables (one per resource type) materialized via dbt, the PySpark bundling pipeline that produces FHIR Bundles for Kafka publication, and the FHIR validator integration that gates publication on US Core 6.1 conformance.
  • Build and maintain FHIR-Repository integration components — the Java/Kotlin egress interceptor that captures client-originated FHIR changes, the Flink loopback consumer that merges those changes into CDM, the bundle consumer that ingests CDM-originated bundles into FHIR-Repository.
  • Implement origin-tag-based loop prevention.
  • Implement Cloud Composer DAGs to orchestrate dbt runs, batch ingestion jobs, maintenance operations (Iceberg compaction, snapshot expiration, orphan file cleanup), and data product refresh schedules.
  • Work within the spec-driven development framework — draft unit specs for new components, work with peers on spec review, generate implementation and tests using Code agents with the spec as primary context, iterate until tests pass, and submit code review packages that include the spec, tests, and implementation together.
  • Implement and monitor data quality checks at every layer — DBT tests for staging and CDM, FHIR validator output for serialization, Iceberg metadata observations for storage health, freshness monitors at the source-to-CDM boundary.
  • Participate in code reviews, on-call rotations, and incident response.
  • Optimize pipeline performance — Flink TaskManager sizing, Iceberg compaction tuning, dbt incremental strategy selection, Starburst cluster scaling decisions.
  • Profile production performance and propose changes when SLOs are at risk.
  • Document the components you build through the SDD framework — every code file references its spec; every spec change is reviewed; every acceptance criterion has a corresponding test.
Required Skills and Qualifications
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field
  • 3+ years of hands-on data engineering experience
  • Strong proficiency in Python and SQL.
  • PySpark and PyFlink familiarity strongly preferred.
  • Some Java or Kotlin exposure useful for FHIR-Repository interceptor work (one or two engineers on the team will lead this; the rest contribute as needed)
  • Hands-on experience with Google Cloud Platform — Cloud Storage, Dataproc, Cloud Composer, Cloud Run or GKE for container workloads, Secret Manager, IAM.
  • Experience with the Dataproc Flink optional component is a strong plus
  • Production experience with dbt — incremental materialization strategies, custom macros, tests, sources, sequencing, and project organization for large model graphs.
  • dbt-trino adapter experience is a plus
  • Production experience with Apache Iceberg — table creation, partitioning, compaction, snapshot operations, schema evolution.
  • Familiarity with reading and writing Iceberg from multiple engines (Spark, Flink, Trino) is valuable
  • Experience with Apache Kafka — producers, consumers, partitioning, consumer-group semantics, retention and compaction, and integration with stream processors.
  • Confluent Cloud experience preferred
  • Experience with streaming data processing — Apache Flink in production preferred; Apache Spark Structured Streaming acceptable as adjacent experience
  • Familiarity with healthcare data standards — at minimum, HL7v2 message structure and FHIR R4 resource shapes.
  • Hands‑on parsing experience for one or both is preferred
  • Experience with version control (Git), branch-based development workflows, pull request reviews, and CI/CD pipelines (GitHub Actions, GitLab CI, or Cloud Build)
  • Comfortable working with AI coding assistants (Code agents, Cursor, Copilot) as collaborators.
  • Strong problem-solving skills, debugging discipline, attention to detail, and ability to operate in an agile team environment
Nice-to-Have Skills
  • FHIR-Repository or HAPI FHIR experience — interceptor authorship, MDM module configuration, FHIR API customization
  • Informatica MDM experience
  • Atlan experience for governance and lineage integration
  • Java or Kotlin proficiency for FHIR-Repository interceptor and HAPI FHIR work
  • Apache Flink production operational experience including stateful jobs, exactly‑once semantics, and savepoint/checkpoint management
  • Experience with FHIR profile validation tools (HL7 FHIR validator, Inferno, or HAPI's validation modules)
  • Experience contributing to or operating spec-driven or contract-first development workflows
  • Experience with Iceberg's Polaris, Nessie, or Snowflake Open Catalog as alternatives to BigLake (useful for portability discussions)

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us! Thank you for your interest in Quantiphi! We are a team that dreams together and collaborates to ensure success. We are currently looking for exceptional individuals who can join us and contribute to a fun, diverse and hybrid work culture. As a part of the Quantiphi family, you will get ample opportunities to learn, grow and interact with colleagues from varied experience and backgrounds around the globe. Quantiphi is an award-winning AI-first digital engineering company driven by the desire to reimagine and realize transformational opportunities at the heart of business. We solve the toughest and most complex business problems with the latest and cutting-edge technologies and to make this happen, we have a vibrant, diverse and talented set of professionals we proudly refer to as the Quantiphi family. Our culture is built on transparency, diversity, integrity, learning and growth. We are committed to provide our teams with an environment that helps them to learn, grow and flourish in their professional as well as personal lives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Quantiphi, Inc. • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Architect - Data Modeller
Architect - Data Modeller

Quantiphi Analytics Solutions Private Limited • Mumbai

On-site
INR 4,000,000 - 5,500,000
Associate Lead Testing
Associate Lead Testing

Quantiphi Analytics Solutions Private Limited • Mumbai

On-site
INR 3,800,000 - 7,000,000
Tech Architect - Data
Tech Architect - Data

Quantiphi Analytics Solutions Private Limited • Mumbai

On-site
INR 2,500,000 - 4,500,000
Senior Data Engineer - India - (Profisee)
Senior Data Engineer - India - (Profisee)

Quantiphi Analytics Solutions Private Limited • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Software Developer (Angular)
Senior Software Developer (Angular)

Quantiphi Analytics Solutions Private Limited • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Data Engineer (SSIS/SSRS)
Senior Data Engineer (SSIS/SSRS)

Quantiphi Analytics Solutions Private Limited • Bengaluru

On-site
INR 2,200,000 - 4,000,000
Research Engineer
Research Engineer

Quantiphi • Mumbai

Hybrid
INR 900,000 - 1,500,000
Hybrid work model
Learning opportunities
Architect – Full stack (React & Java)
Architect – Full stack (React & Java)

Quantiphi Analytics Solutions Private Limited • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Associate Architect - Data
Associate Architect - Data

Quantiphi Analytics Solutions Private Limited • Mumbai

On-site
INR 2,000,000 - 4,200,000