Lead Data Engineer

US Pharmacopeia

Hyderabad

On-site

INR 2,400,000 - 3,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Company-paid time off
Health care options
Retirement savings plans

Job summary

USP’s Digital Product Engineering team seeks a Data Engineer to build robust data pipelines and manage large-scale data infrastructure, enabling advanced analytics for patient safety and health worldwide.

The ideal candidate will drive data architecture, cloud technologies, and scalable engineering solutions with strong expertise in Spark, SQL, and BI integrations, collaborating across business, product, and analytics teams.

Qualifications

  • Bachelor's degree in a relevant field (e.g., Engineering, Analytics, Data Science, Computer Science, Statistics) or equivalent experience.
  • 7+ years of experience with big-data technologies such as Python, PySpark, and SQL for processing structured, semi-structured, and unstructured data.
  • Strong experience with AWS data services including Redshift, S3, Glue, Lambda, EventBridge, and Postgres; experience with Azure/GCP equivalents acceptable.
  • Experience building batch, micro-batch, and streaming pipelines using Lambda/Kappa architectures.
  • Experience designing and delivering enterprise-scale data platforms (lakehouse, warehouse, lake, marts).
  • Solid understanding and hands-on implementation of data-modeling techniques such as Data Vault 2.0, dimensional modeling, knowledge graphs, and OBT approaches; certification preferred.
  • Experience with medallion architecture and metadata-driven pipeline frameworks.
  • Deep knowledge of data governance frameworks, including data discovery, quality, security; experience with DQ tools such as Great Expectations, Pydantic, etc.
  • Strong SQL and programming skills for data transformation, modeling, and analysis.
  • Hands-on experience building and maintaining complex ETL/ELT pipelines and day-to-day data operations.
  • Experience with workflow orchestration tools such as Airflow (2+ years or equivalent).
  • Understanding of streaming/event-driven architectures and modern data processing patterns.
  • Proficiency in dashboards and visualization with Tableau, Power BI, or equivalent.
  • Experience with Agile (SAFe) methodologies, CI/CD pipelines, and modern deployment practices for data platforms.
  • Exposure to AI/ML concepts, familiarity with generative-AI patterns (e.g., RAG, chunking).

Responsibilities

  • Design and implement scalable data collection, storage, and processing pipelines to support enterprise-wide data needs.
  • Implement and maintain data governance frameworks and data quality checks within data pipelines to ensure compliance and reliability.
  • Build and optimize data models and data marts to support self-service analytics and reporting tools such as Tableau, Looker, and Power BI.
  • Partner with data scientists to operationalize models by integrating them into production-grade pipelines, ensuring scalability, performance, and maintainability.
  • Collaborate with cross-functional stakeholders (business, product, analytics, engineering teams) to translate business requirements into scalable data solutions and prioritize data initiatives.
  • Provide technical leadership and architectural guidance for data platform design, ensuring alignment with enterprise standards and long-term data strategy.
  • Lead design reviews, code reviews, and data architecture discussions to ensure best practices, reusability, and high-quality deliverables.
  • Collaborate with platform, DevOps, and security teams to ensure secure, cost-effective, and scalable data infrastructure.
  • Influence the data roadmap and strategy by identifying opportunities for platform enhancement, automation, and cost optimization.

Skills

Python
PySpark
SQL
AWS
Airflow
ETL/ELT
Data modeling
Tableau/Power BI
Data governance
CI/CD

Education

Bachelor's degree

Tools

Redshift
S3
Glue
Lambda
EventBridge

Job description

About USP

The U.S. Pharmacopeial Convention (USP) is an independent scientific organization that collaborates with the world’s top experts in health and science to develop quality standards for medicines, dietary supplements, and food ingredients. USP’s fundamental belief that Equity = Excellence manifests in our core value of Passion for Quality, and we have more than 1,100 professionals across five global locations, working to strengthen the supply of safe, quality medicines and supplements worldwide.

Job Overview

The Digital Product Engineering team at USP seeks a Data Engineer who will build robust data pipelines, manage large-­scale data infrastructure, and enable advanced analytics capabilities. This role supports projects aligned with our mission to protect patient safety and improve health worldwide. The ideal candidate brings passion for data architecture, cloud technologies, and scalable engineering solutions that drive innovation and impact.

Responsibilities
  • Design and implement scalable data collection, storage, and processing pipelines to support enterprise-wide data needs.
  • Implement and maintain data governance frameworks and data quality checks within data pipelines to ensure compliance and reliability.
  • Build and optimize data models and data marts to support self‑service analytics and reporting tools such as Tableau, Looker, and Power BI.
  • Partner with data scientists to operationalize models by integrating them into production‑grade pipelines, ensuring scalability, performance, and maintainability.
  • Collaborate with cross‑functional stakeholders (business, product, analytics, engineering teams) to translate business requirements into scalable data solutions and prioritize data initiatives.
  • Provide technical leadership and architectural guidance for data platform design, ensuring alignment with enterprise standards and long‑term data strategy.
  • Lead design reviews, code reviews, and data architecture discussions to ensure best practices, reusability, and high‑quality deliverables.
  • Collaborate with platform, DevOps, and security teams to ensure secure, cost‑effective, and scalable data infrastructure.
  • Influence the data roadmap and strategy by identifying opportunities for platform enhancement, automation, and cost optimization.
Qualifications

Education

  • Bachelor’s degree in a relevant field (e.g., Engineering, Analytics, Data Science, Computer Science, Statistics) or equivalent experience.

Experience

  • 7+ years of experience with big‑data technologies such as Python, PySpark, and SQL for processing structured, semi‑structured, and unstructured data.
  • Strong experience with AWS data services including Redshift, S3, Glue, Lambda, EventBridge, and Postgres; experience with Azure/GCP equivalents acceptable.
  • Experience building batch, micro‑batch, and streaming pipelines using Lambda/Kappa architectures.
  • Experience designing and delivering enterprise‑scale data platforms (lakehouse, warehouse, lake, marts).
  • Solid understanding and hands‑on implementation of data‑modeling techniques such as Data Vault 2.0, dimensional modeling, knowledge graphs, and OBT approaches; certification preferred.
  • Experience with medallion architecture and metadata‑driven pipeline frameworks.
  • Deep knowledge of data governance frameworks, including data discovery, quality, security; experience with DQ tools such as Great Expectations, Pydantic, etc.
  • Strong SQL and programming skills for data transformation, modeling, and analysis.
  • Hands‑on experience building and maintaining complex ETL/ELT pipelines and day‑to‑day data operations.
  • Experience with workflow orchestration tools such as Airflow (2+ years or equivalent).
  • Understanding of streaming/event‑driven architectures and modern data processing patterns.
  • Proficiency in dashboards and visualization with Tableau, Power BI, or equivalent.
  • Experience with Agile (SAFe) methodologies, CI/CD pipelines, and modern deployment practices for data platforms.
  • Exposure to AI/ML concepts, familiarity with generative‑AI patterns (e.g., RAG, chunking).
Preferred Qualifications
  • Experience with scientific chemistry nomenclature or prior work in life sciences, chemistry, or hard sciences.
  • Experience with pharmaceutical datasets and nomenclature.
  • Experience developing machine‑learning and deep‑learning models; familiarity with MLOps and deploying ML models in production.
  • Ability to explain complex technical issues to a non‑technical audience.
Benefits

USP provides comprehensive benefits, including company‑paid time off, health care options, retirement savings plans, and other protections for you and your family.

Other Information

Job Category: Information Technology
Job Type: Full‑Time

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

Initial Therapeutics, Inc. • India

On-site
INR 1,800,000 - 3,400,000
Healthcare options
Retirement savings plan
Company-paid time off
Lead Data Engineer
Lead Data Engineer

BioSpace • Chittoor

On-site
INR 3,000,000 - 4,800,000
Healthcare
Retirement plan
Paid time off
Data Scientist, Digital Products
Data Scientist, Digital Products

BioSpace • Chittoor

On-site
INR 1,200,000 - 2,000,000
Data Scientist, Digital Products
Data Scientist, Digital Products

Initial Therapeutics, Inc. • India

On-site
INR 800,000 - 1,200,000
Company-paid time off
Comprehensive healthcare options
Retirement savings plans
Data Scientist, Digital Products
Data Scientist, Digital Products

US Pharmacopeia • Hyderabad

On-site
INR 600,000 - 1,200,000
Company-paid time off
Comprehensive healthcare options
Retirement savings
Digital Products Manager
Digital Products Manager

The U.S. Pharmacopeial Convention (USP) • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Digital Products Manager
Digital Products Manager

US Pharmacopeia • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Lead Big Data Engineer
Lead Big Data Engineer

S&P Global, Inc. • Rangareddy

On-site
INR 1,800,000 - 2,600,000
Health & Wellness
Flexible downtime
Continual learning
+3
Data Analyst, Sr
Data Analyst, Sr

Transformcap • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Data Engineer - Pyspark, Databricks, Snowflake, Azure Cloud
Data Engineer - Pyspark, Databricks, Snowflake, Azure Cloud

UnitedHealth Group • Hyderabad

On-site
Confidential