Sr Data Engineer

Integrichain

Town of Philadelphia (NY)

On-site

USD 130,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical benefits
Flexible PTO
Learning and development

Job summary

IntegriChain, the data backbone for life sciences, seeks a senior data engineer to lead data integration, MDM, and analytics platform design.

You will design Snowflake data models, develop scalable ELT pipelines using dbt, and collaborate with product and data science teams to deliver trusted data assets for business decisions.

Qualifications

  • 10+ years in data engineering, analytics engineering, or data platform development in production settings.

Responsibilities

  • Define and mature enterprise data integration, data consolidation, MDM integration, and data platform design patterns across the organization.

Skills

Snowflake
SQL/PLSQL
Python
ETL/ELT
dbt
MDM concepts
Data modeling
APIs / Files integration
Data quality / observability
Cross-functional collaboration
Airflow / orchestration
Terraform / IaC

Tools

Airflow
Dagster
dbt Cloud
Terraform
CI/CD

Job description

Requirements
Must have:
  • We need 10+ years of experience in data engineering, database engineering, analytics engineering, or data platform development in production settings.
  • We need strong hands-on expertise with Snowflake, including architecture, performance tuning, security design, cost optimization, and cost tracking.
  • We need a solid understanding of Snowflake design patterns for analytical workloads, high-volume processing, data sharing, and multi-environment deployments.
  • We need practical experience with ETL/ELT tools; dbt experience is strongly preferred.
  • We need strong SQL and PL/SQL-style development experience, including complex transformations, stored procedures, tuning, and large-scale processing.
  • We need Python experience for data automation, API integration, file handling, data validation, metadata processing, or operational tooling.
  • We need experience designing enterprise data models, curated layers, semantic layers, and reusable data products.
  • We need experience with data integration across enterprise applications, APIs, files, cloud storage, operational systems, MDM platforms, and analytics platforms.
  • We need a working knowledge of Master Data Management concepts such as golden records, crosswalks/XREFs, match/merge, survivorship, hierarchies, entity relationships, stewardship, and data quality.
  • We need experience partnering with MDM, Product, or business teams to translate mastering needs into mappings, transformation logic, validations, and downstream consumption patterns.
  • We need the ability to work directly with cross-functional stakeholders to gather requirements, explain design tradeoffs, and build alignment.
  • We need experience implementing data quality, lineage, auditability, observability, and operational monitoring in data pipelines.
  • We need the confidence to operate as a hands-on senior individual contributor while also influencing strategy and engineering standards.
  • Preferred: experience with Reltio MDM, including inbound loads, outbound exports, APIs, crosswalks, match/merge outputs, survivorship outputs, and troubleshooting.
  • Preferred: experience in life sciences, healthcare, pharma commercialization, HCO/HCP mastering, patient data, channel data, customer master, or commercial data platforms.
  • Preferred: experience with life sciences reference and commercial datasets such as HIN, DEA, NPI, NCPDP, 340B/PHS, 844, 852, 867, chargebacks, gross-to-net, government pricing, PBR, or UBR.
  • Preferred: experience with orchestration frameworks such as Airflow, Dagster, dbt Cloud jobs, cloud-native schedulers, or similar tools.
  • Preferred: experience with cloud platforms and storage patterns, especially Azure or AWS object storage integrated with Snowflake.
  • Preferred: exposure to AI-ready data architecture, feature stores, ML datasets, semantic models, or AI/ML pipeline enablement.
  • Preferred: experience with Terraform, CI/CD, Git-based development, and infrastructure-as-code practices.
  • Preferred: Snowflake SnowPro, Reltio, dbt, or comparable cloud/data engineering certifications.
Responsibilities:
  • We help define and mature enterprise data integration, data consolidation, MDM integration, and data platform design patterns across our organization.
  • We design, build, optimize, and operate Snowflake data models, pipelines, stored procedures, and high-volume processing patterns.
  • We partner with MDM and Product teams to support HCO Master data ingestion, outbound extracts, cross-reference data, golden record consumption, survivorship outputs, and downstream publishing.
  • We collaborate with Product, Engineering, MDM, Data Science, DevOps, Security, and business stakeholders to align data solutions with enterprise priorities.
  • We use dbt or similar ELT tooling to create reliable, maintainable, testable, and observable pipelines.
  • We drive Snowflake performance tuning, warehouse sizing, workload management, cost tracking, and cost optimization.
  • We partner with Data Science leadership to rationalize and consolidate the enterprise data landscape across products, platforms, and acquired capabilities.
  • We define reusable integration patterns for batch, micro-batch, near-real-time, and application-to-application exchange.
  • We design scalable patterns for ingesting, transforming, mastering, and publishing data across operational and analytical use cases.
  • We establish standards for data contracts, schema evolution, data quality, lineage, and data ownership.
  • We build pipelines that load source data into Reltio MDM and extract mastered outputs for downstream Snowflake, analytics, AI, and operational uses.
  • We translate HCO mastering requirements into pipeline, mapping, validation, reconciliation, and publishing patterns.
  • We work with Reltio APIs, exports, crosswalks/XREFs, event-based integrations, and bulk load/extract mechanisms.
  • We engineer integration patterns for HCO Master data, including party/entity, address, identifier, hierarchy, relationship, match/merge, survivorship, and golden record outputs.
  • We support source ingestion and reference data integration for datasets such as HIN, DEA, NPI, NCPDP, 340B/PHS, channel outlet data, customer/account data, and other master/reference sources.
  • We develop validation and reconciliation processes across source data, Reltio mastered data, Snowflake curated data, and downstream consumption layers.
  • We operationalize MDM outputs for business-facing data products, semantic models, reporting tables, APIs, and AI-ready datasets.
  • We design Snowflake database, schema, table, view, and semantic-layer patterns that support performance, governance, and maintainability.
  • We optimize Snowflake workloads using clustering, micro-partition awareness, warehouse sizing, query profiling, caching behavior, and workload isolation.
  • We implement cost tracking and optimization practices, including warehouse utilization monitoring, inefficient query detection, and cost allocation.
  • We build scalable SQL and Snowflake stored procedure logic for large-volume data processing and analytics.
  • We apply secure Snowflake design patterns including RBAC, masking, access isolation, auditing, and environment separation.
  • We design, build, and maintain dependable ELT pipelines using dbt or comparable modern transformation tooling.
  • We develop Python automation for API integration, file processing, metadata management, validation, orchestration support, and operational tooling.
  • We create modular, tested, and reusable transformation models for raw, curated, mastered, and business-ready layers.
  • We implement automated data quality checks, freshness checks, reconciliation, logging, and exception handling.
  • We build orchestration-ready pipelines that support dependency management, restartability, incremental loads, and operational monitoring.
  • We collaborate with DevOps/SRE teams on CI/CD, deployment automation, environment promotion, and operational runbooks.
  • We lead logical and physical data modeling for enterprise analytical, operational, MDM, and AI-ready datasets.
  • We design models that balance normalization, dimensional modeling, medallion/lakehouse concepts, and application-specific consumption needs.
  • We create denormalized reporting and semantic-model-ready structures that simplify business use and reduce ambiguity for AI/LLM use cases.
  • We process and optimize large data volumes in Snowflake using efficient SQL, procedural logic, Snowflake Scripting, and performance-aware design.
  • We create reusable patterns for historical tracking, snapshots, audit columns, data versioning, and lifecycle management.
  • We ensure data models support BI, AI/ML, semantic models, data apps, MDM Explorer/Entity 360 use cases, and enterprise reporting.
Company

We are IntegriChain, the data and application backbone for market access departments at life sciences manufacturers. Our ICyte platform supports commercial and government payer contracting, patient services, and distribution by uniting the financial, operational, and commercial data needed to accelerate therapy access in specialty and precision medicine. We are headquartered in Philadelphia, Pennsylvania, with additional offices in Ambler, Pennsylvania; Pune, India; and Medellín, Colombia. We offer a mission-driven environment focused on improving patients lives, along with excellent and affordable medical benefits, flexible paid time off, and robust learning and development opportunities, including more than 700 free development courses for employees. We are committed to equal opportunity, inclusion, and a workplace free from discrimination and harassment, and applicants for U.S.-based roles must have valid work authorization without future visa sponsorship needs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Manager - Data Science (MDM / Net Revenue)
Product Manager - Data Science (MDM / Net Revenue)

IntegriChain • Philadelphia

On-site
USD 100,000 - 130,000
Flexible Paid Time Off
Learning & Development opportunities
Product Analyst, Conversational AI
Product Analyst, Conversational AI

IntegriChain • Philadelphia

On-site
USD 110,000 - 160,000
Flexible PTO
Medical benefits
Learning & Development
+1
Product Analyst, Agentic AI
Product Analyst, Agentic AI

IntegriChain • Philadelphia

On-site
USD 95,000 - 130,000
Medical benefits
Flexible Paid Time Off
Learning & Development opportunities
Technical Product Analyst
Technical Product Analyst

IntegriChain • Philadelphia

Hybrid
USD 90,000 - 130,000
Medical benefits
Student Loan Reimbursement
Flexible Paid Time Off
+3
Technical Business Analyst
Technical Business Analyst

IntegriChain Incorporated. • Philadelphia

Hybrid
USD 80,000 - 100,000
Excellent medical benefits
Student Loan Reimbursement
Flexible Paid Time Off
+3
Automation & Process Improvement Engineer
Automation & Process Improvement Engineer

IntegriChain • Pennsylvania

Hybrid
USD 110,000 - 170,000
Student Loan Reimbursement
Flexible Paid Time Off
Paid Parental Leave
+2
Senior BI Developer
Senior BI Developer

IntegriChain • Philadelphia

On-site
USD 110,000 - 150,000
Medical benefits
Flexible PTO
Learning & development
Customer Engagement Manager (West Coast)
Customer Engagement Manager (West Coast)

IntegriChain Incorporated • Northern (KY)

Hybrid
USD 90,000 - 120,000
Student Loan Reimbursement
Flexible Paid Time Off
Paid Parental Leave
+2
Product Analyst, Agentic AI
Product Analyst, Agentic AI

Vibehackers • United States

On-site
USD 85,000 - 120,000
Medical benefits
Flexible PTO
Learning & development (700+ courses)
+1
Senior BI Developer
Senior BI Developer

IntegriChain Incorporated • Philadelphia

Hybrid
USD 120,000 - 150,000
Flexible Paid Time Off (PTO)
700+ development courses
Excellent medical benefits