About USP
The U.S. Pharmacopeial Convention (USP) is an independent scientific organization that collaborates with the world’s top experts in health and science to develop quality standards for medicines, dietary supplements, and food ingredients. USP’s fundamental belief that Equity = Excellence manifests in our core value of Passion for Quality, and we have more than 1,100 professionals across five global locations, working to strengthen the supply of safe, quality medicines and supplements worldwide.
Job Overview
The Digital Product Engineering team at USP seeks a Data Engineer who will build robust data pipelines, manage large-scale data infrastructure, and enable advanced analytics capabilities. This role supports projects aligned with our mission to protect patient safety and improve health worldwide. The ideal candidate brings passion for data architecture, cloud technologies, and scalable engineering solutions that drive innovation and impact.
Responsibilities
- Design and implement scalable data collection, storage, and processing pipelines to support enterprise-wide data needs.
- Implement and maintain data governance frameworks and data quality checks within data pipelines to ensure compliance and reliability.
- Build and optimize data models and data marts to support self‑service analytics and reporting tools such as Tableau, Looker, and Power BI.
- Partner with data scientists to operationalize models by integrating them into production‑grade pipelines, ensuring scalability, performance, and maintainability.
- Collaborate with cross‑functional stakeholders (business, product, analytics, engineering teams) to translate business requirements into scalable data solutions and prioritize data initiatives.
- Provide technical leadership and architectural guidance for data platform design, ensuring alignment with enterprise standards and long‑term data strategy.
- Lead design reviews, code reviews, and data architecture discussions to ensure best practices, reusability, and high‑quality deliverables.
- Collaborate with platform, DevOps, and security teams to ensure secure, cost‑effective, and scalable data infrastructure.
- Influence the data roadmap and strategy by identifying opportunities for platform enhancement, automation, and cost optimization.
Qualifications
Education
- Bachelor’s degree in a relevant field (e.g., Engineering, Analytics, Data Science, Computer Science, Statistics) or equivalent experience.
Experience
- 7+ years of experience with big‑data technologies such as Python, PySpark, and SQL for processing structured, semi‑structured, and unstructured data.
- Strong experience with AWS data services including Redshift, S3, Glue, Lambda, EventBridge, and Postgres; experience with Azure/GCP equivalents acceptable.
- Experience building batch, micro‑batch, and streaming pipelines using Lambda/Kappa architectures.
- Experience designing and delivering enterprise‑scale data platforms (lakehouse, warehouse, lake, marts).
- Solid understanding and hands‑on implementation of data‑modeling techniques such as Data Vault 2.0, dimensional modeling, knowledge graphs, and OBT approaches; certification preferred.
- Experience with medallion architecture and metadata‑driven pipeline frameworks.
- Deep knowledge of data governance frameworks, including data discovery, quality, security; experience with DQ tools such as Great Expectations, Pydantic, etc.
- Strong SQL and programming skills for data transformation, modeling, and analysis.
- Hands‑on experience building and maintaining complex ETL/ELT pipelines and day‑to‑day data operations.
- Experience with workflow orchestration tools such as Airflow (2+ years or equivalent).
- Understanding of streaming/event‑driven architectures and modern data processing patterns.
- Proficiency in dashboards and visualization with Tableau, Power BI, or equivalent.
- Experience with Agile (SAFe) methodologies, CI/CD pipelines, and modern deployment practices for data platforms.
- Exposure to AI/ML concepts, familiarity with generative‑AI patterns (e.g., RAG, chunking).
Preferred Qualifications
- Experience with scientific chemistry nomenclature or prior work in life sciences, chemistry, or hard sciences.
- Experience with pharmaceutical datasets and nomenclature.
- Experience developing machine‑learning and deep‑learning models; familiarity with MLOps and deploying ML models in production.
- Ability to explain complex technical issues to a non‑technical audience.
Benefits
USP provides comprehensive benefits, including company‑paid time off, health care options, retirement savings plans, and other protections for you and your family.
Other Information
Job Category: Information Technology
Job Type: Full‑Time