The Professional, Data Engineer designs, builds, and maintains the data pipelines, integrations, and platforms that form the foundation of analytics, reporting, and machine learning across DePuy Synthes. Sitting within the Data & AI organization, this role is responsible for ingesting, processing, cleaning, and transforming structured and unstructured data from diverse source systems into reliable, high-quality, and accessible data assets. As an established individual contributor working with minimal supervision, the engineer partners with data architects, data modelers, data scientists, and business stakeholders to define data requirements, optimize data/information flow, and increase the efficiency and automation of data processes. The role directly supports the organization's overall Data Analytics & Computational Sciences strategy and the build-out of enterprise data capabilities.
- Design, build, and maintain scalable, reliable ETL/ELT pipelines to ingest, transform, and integrate data from multiple source systems.
- Architect and build data pipelines and aggregation systems to deliver quality real-time and batch analytical datasets.
- Develop and optimize data models and structures that enable efficient storage, retrieval, and downstream analytics and machine learning.
- Process moderately complex structured and unstructured datasetschunking, parsing, wrangling, cleaning, and segregating data—to create fit-for-purpose data structures.
- Perform data collection, processing, cleaning, analysis, modeling, and visualization to support business questions and improve decision-making.
- Build and manage workflow orchestration using tools such as Apache Airflow or Azure Data Factory (ADF) to automate and schedule data pipelines.
- Develop distributed data processing solutions using Spark for large-scale data transformation and analysis.
- Optimize pipeline performance, cost, and reliability, identifying and resolving bottlenecks and data integration failures.
- Ensure data integrity, quality, security, and compliance throughout the data lifecycle, adhering to data governance and privacy standards.
- Implement data quality checks, validation rules, and monitoring to maintain clean and healthy data across upstream and downstream channels.
- Collaborate with cross-functional teams to provide a data-driven perspective on business problems and translate requirements into technical solutions.
- Develop custom programs and reusable components to address unique data engineering challenges and deploy data solutions and platforms.
- Prepare and maintain technical documentation, including data flow diagrams, pipeline specifications, and design decisions.
What you'll bring
Education:
- Required: Bachelor's degree in Computer Science, Information Systems, Data Engineering, Engineering, or a related quantitative/technical field.
- Preferred: Master's degree in Computer Science, Data Engineering, Information Management, or a related discipline.
Experience and Skills
Required
- 3+ years of experience in data engineering, ETL/ELT development, or a related data-focused role.
- Advanced SQL skills, including query optimization and performance tuning.
- Strong proficiency in Python for data processing, automation, and pipeline development.
- Hands-on experience designing and building ETL/ELT pipelines and data integration workflows.
- Experience with distributed data processing using Apache Spark.
- Experience with cloud data platforms (Azure and/or AWS).
- Hands-on experience with Databricks (Delta Lake, Databricks SQL, Spark) for building scalable data pipelines.
- Experience with workflow orchestration tools (e.g., Apache Airflow, Azure Data Factory).
- Solid understanding of data modeling concepts and data warehousing principles.
Preferred
- Experience in a regulated industry (MedTech, Pharmaceutical, or Healthcare) with familiarity in associated data governance and compliance requirements.
- Experience with cloud data warehouse/lakehouse platforms (e.g., Databricks, Snowflake, Azure Synapse, Redshift, BigQuery).
- Familiarity with data cataloging, lineage, and master data management tools.
- Exposure to streaming/real-time data processing frameworks.
- Experience working in Agile environments and cross-functional product teams.