The Patient Data Aggregator is responsible for the secure collection, standardization, and aggregation of sensitive patient data across multiple systems to support clinical, operational, and research needs. This role requires a strong understanding of data privacy protocols, including de‑identification methodologies such as tokenization and expert determination, and must operate within a HISEC‑compliant (High Security) environment. The candidate will collaborate with cross‑functional teams to ensure data is managed in a secure, compliant, and usable manner.
Responsibilities
- Collect, clean, and aggregate patient‑level data from Electronic Health Records (EHRs), lab systems, and external databases.
- Apply tokenization techniques to replace direct identifiers while preserving data integrity for linkage and analysis.
- Perform or support expert determination de‑identification processes to evaluate and minimize re‑identification risk in datasets.
- Ensure all activities comply with HIPAA, and institutional privacy policies, particularly within HISEC environments.
- Manage and maintain secure data pipelines and repositories for sensitive patient information.
- Assist in the development of standard operating procedures (SOPs) for secure data aggregation and handling.
- Support the generation of anonymized datasets for research, analytics, or external data sharing initiatives.
- Monitor data quality, flag inconsistencies, and ensure proper documentation of data provenance and transformation.
- Provide support for audits, security assessments, and internal data governance reviews.
- Design and implement end‑to‑end Data aggregation solutions tailored to business requirements using Python and Pyspark.
- Synthesizes findings, develops recommendations, and communicates results to clients and internal teams.
- Assumes project management responsibilities for Data aggregation implementation on each project with minimal supervision, including managing onshore/client communication, leading meetings, and drafting agendas.
- Manage multiple projects simultaneously while meeting deadlines.
Qualifications
- Engineering/Master’s Degree from a Tier‑I/Tier‑II institution/university in Computer Science or relevant concentration, with evidence of strong academic performance.
- 6+ years of relevant consulting‑industry experience.
- Deep understanding of data management best practices, data modeling and data analytics.
- Strong understanding of U.S. pharmaceutical datasets and their applications.
- Experience with tokenization tools/methods and expert determination for de‑identification.
- Logical thinking and problem‑solving skills along with an ability to collaborate.
- Familiarity with HISEC standards or other high‑security frameworks.
- Strong programming skills in Python and working knowledge of PySpark.
- Proficiency in Microsoft Office products, data analysis tools (e.g., Snowflake, Redshift) and job orchestration (Airflow, Databricks workflows).
- Strong understanding of HIPAA, GDPR, and related data privacy regulations.
- Excellent verbal and written communication skills, with the ability to convey complex information clearly to diverse audiences.
- Strong organizational abilities and effective time management skills.
- Ability to thrive in an international matrix environment and willingness to support US clients during US working hours with 3‑4 hours of overlap.