Principal Data Engineer (RWE)

Jobgether

India

On-site

INR 6,000,000 - 9,000,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Flexible working environment
Large-scale healthcare datasets
Cross-functional collaboration

Job summary

Jobgether is seeking a Principal Data Engineer to build scalable data products for Real World Evidence, observational research, and dashboards. You will transform heterogeneous healthcare datasets into reusable data models and OMOP formats where appropriate, while ensuring FAIR data principles and AI-ready pipelines.

You will collaborate with epidemiologists, statisticians, market access specialists, and IT teams to translate business requirements into robust data structures and scalable ETL/ELT

Qualifications

  • Deep understanding of OMOP CDM v5.4/v6 and extensions.
  • Strong expertise with Databricks, PySpark, Spark SQL, SQL, and Delta Lake.
  • Experience with large-scale healthcare datasets and distributed processing.
  • Knowledge of SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT.
  • Ability to translate business needs into technical data structures.
  • Experience building AI-ready datasets and FAIR-compliant pipelines.
  • Excellent stakeholder management and cross-functional collaboration.
  • Nice-to-have: R, sparklyR, R Shiny, OHDSI tools, and Azure Data Platform services.

Responsibilities

  • Develop automated data processes for patient-level data products supporting dashboards, reports, studies, and other business needs.
  • Transform raw datasets into usable, structured data products for RWE studies and dashboards.
  • Convert heterogeneous datasets into reusable data models for observational research and epidemiology.
  • Convert datasets into OMOP format where appropriate and identify non-standard data usability.
  • Build FAIR data pipelines and semantic data frameworks to improve interoperability.
  • Develop AI-ready datasets and data products for generative AI use cases.
  • Engage with stakeholders to translate business requirements into technical data structures.
  • Collaborate with the RWE programming team to support data structures for study outputs.
  • Liaise with IT to ensure inbound datasets are fit for analysis.
  • Collaborate with data/vendor teams, e.g., Databricks, as needed.
  • Maintain documentation covering data flows, schemas, pipelines, and processes.
  • Design testing, validation, and monitoring to ensure data quality.
  • Troubleshoot ETL data loading and transformation issues.
  • Collaborate with team members to share workload as needed.

Skills

Databricks
PySpark
Spark SQL
SQL
Delta Lake
Data Modeling
FAIR data
RWE
ETL/ELT pipelines
Cloud platforms

Tools

OHDSI tools
Azure Data Platform
Power BI

Job description

As a Principal Data Engineer, you will play a key role in building scalable data products that support Real World Evidence (RWE), observational research, epidemiology studies, dashboards, and other analytical outputs. You will transform complex, heterogeneous healthcare datasets into reusable, high-quality data models and pipelines. The role combines advanced data engineering with deep healthcare data expertise, including OMOP standards and clinical terminologies. You will collaborate closely with epidemiologists, statisticians, health economists, market access specialists, IT teams, and other technical stakeholders. You will also contribute to FAIR data principles and the development of AI-ready datasets for emerging generative AI use cases. This is an opportunity to work with large-scale patient-level data in a collaborative, globally connected data engineering environment.

Accountabilities:
  • Develop automated data processes for the ongoing generation of patient-level data products supporting dashboards, reports, studies, and other business needs.
  • Transform raw datasets received from data vendors and partners into usable, structured data products for RWE studies, dashboards, and analytical outputs.
  • Convert heterogeneous healthcare datasets into reusable data models that support observational research and epidemiology studies.
  • Convert bespoke datasets, such as biomarkers and mutation data, into OMOP format where appropriate, while identifying residual data that cannot be standardized and determining how it can still support analysis.
  • Build FAIR (Findable, Accessible, Interoperable, Reusable) data pipelines and semantic data engineering frameworks that improve healthcare data discoverability and interoperability.
  • Develop AI-ready datasets and data products capable of supporting generative AI and other advanced analytics use cases.
  • Engage with epidemiologists, statisticians, market access specialists, health economists, and other stakeholders to understand, scope, document, and translate business requirements into actionable technical data structures.
  • Collaborate with the RWE programming team to develop data structures required for study outputs and provide technical support where data engineering expertise adds value.
  • Liaise with IT teams to ensure inbound datasets from data partners are fit for their intended analytical purposes.
  • Collaborate with technical teams from data and analytics software vendors, including Databricks, when required.
  • Maintain clear documentation covering data flows, schemas, pipelines, and processes to support onboarding, troubleshooting, and auditing.
  • Design and implement comprehensive testing, validation, and monitoring approaches to ensure the accuracy, reliability, and quality of data products.
  • Troubleshoot issues related to data loading, extraction, transformation, and ETL processes.
  • Collaborate with other members of the Data Engineering team, providing support and taking on additional workload when needed.
Requirements
  • Strong understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts and their application to healthcare analytics.
  • Ability to assess business requirements and recommend appropriate real-world healthcare datasets for analytical use cases.
  • Deep understanding of healthcare data models, healthcare data ecosystems, and patient-level datasets.
  • Strong expertise in OMOP Common Data Model v5.4 and v6, including extensions.
  • Knowledge of healthcare standards and terminologies such as SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT.
  • Strong experience building scalable ETL/ELT pipelines using Databricks, PySpark, Spark SQL, SQL, and Delta Lake.
  • Experience working with large-scale healthcare and patient-level datasets and distributed data processing frameworks.
  • Strong understanding of Semantic Data Engineering principles and experience developing FAIR-compliant data pipelines.
  • Experience with cloud-based data platforms and large-scale distributed processing environments.
  • Strong Power BI development and data modeling capabilities, with the ability to create reusable analytical datasets for dashboards and studies.
  • Experience designing AI-ready datasets and analytics data products.
  • Strong data profiling, validation, monitoring, and automated data quality framework experience.
  • Understanding of healthcare data quality assessment methodologies.
  • Excellent stakeholder management, communication, and collaboration skills, with the ability to translate complex business needs into effective technical solutions.
  • Experience working with cross-functional and globally distributed teams.
  • Exposure to one or more therapeutic areas such as Oncology, Respiratory, Immunology and Inflammation, or Infectious Diseases.
  • Nice-to-have: working knowledge of R, sparklyR, R Shiny, observational research methodologies, OHDSI tools, and Azure Data Platform services.
Benefits
  • India-based opportunity with a flexible working environment.
  • Opportunity to work with large-scale healthcare and patient-level datasets.
  • Exposure to Real World Data, Real World Evidence, observational research, epidemiology, and healthcare analytics.
  • Opportunity to contribute to FAIR data engineering and AI-ready data initiatives.
  • Collaborative environment with cross-functional and globally distributed teams.
  • Opportunities to work across data engineering, analytics, healthcare standards, and emerging AI use cases.
  • Professional development through collaboration with experienced data engineering, programming, analytics, and healthcare specialists.
  • Inclusive workplace culture that values diversity, integrity, honesty, and respect.

#LI-CL1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Data Engineer (RWE)
Principal Data Engineer (RWE)

Veramed • India

On-site
INR 1,200,000 - 2,000,000
Principal Data Engineer
Principal Data Engineer

Takeda • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Data Engineer
Data Engineer

Excelra • Hyderabad, Bengaluru, Delhi

On-site
INR 1,200,000 - 1,800,000
Principal Data Engineer RWE RWD
Principal Data Engineer RWE RWD

Veramed • India

On-site
INR 700,000 - 1,100,000
Senior Manager Research Analytics
Senior Manager Research Analytics

Putnam • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineering Analyst
Senior Data Engineering Analyst

Optum India • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Manager Data Platform Engineer
Manager Data Platform Engineer

HealthEdge • Bengaluru

On-site
INR 4,200,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

R1 RCM • Dadri

On-site
INR 1,200,000 - 1,800,000
Competitive benefits package
Opportunities for growth and learning
Impactful work in communities
Consulting Data Engineer
Consulting Data Engineer

Elsevier • Delhi

On-site
INR 5,000,000 - 6,000,000
Comprehensive Health Insurance
Flexible Working Arrangements
Employee Assistance Program
+2
Senior Associate - RWE Obesity Statistical Programming
Senior Associate - RWE Obesity Statistical Programming

Amgen • Hyderabad

On-site
INR 1,200,000 - 1,600,000