Data Engineer

Jobtailor

Lebanon (Lebanon County)

On-site

USD 110,000 - 170,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lilly is seeking a data engineer to design and build data pipelines for ingestion, transformation, and integration across AWS and Azure.

You will work with IT/OT systems and collaborate with Operations, Quality, Lab, and Engineering to deliver robust data products while ensuring GxP compliance.

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Engineering, or related field.
  • 3+ years in data modeling, ETL/ELT, ontology development, semantic graphs, or relational schema design.
  • 1+ year in pharma GxP or scientific environment.
  • Experience with AWS and Azure cloud platforms.
  • Experience with AI/ML/LLM concepts and tools.
  • 3+ years designing large-scale data models for functional, operational, and analytical environments.
  • Proficiency in SQL and data modeling.
  • Experience with ER/Studio, Erwin, or TOAD.
  • Experience with streaming, Industrial IoT, MQTT, Kafka.
  • Understanding of data lakehouses, data warehousing, and big data concepts.
  • Security models and development on large datasets.
  • GxP compliance and Computer System Validation knowledge.

Responsibilities

  • Design and build data pipelines for ingestion, transformation, and integration.
  • Integrate IT and OT source systems with AWS and Azure cloud data lakehouse architecture.
  • Capture, ingest, integrate, contextualize, harmonize, and deliver data as reusable data domains and data products.
  • Engage with Operations, Quality, Lab, and Engineering stakeholders to understand data needs.
  • Translate operational needs into technical designs and explain decisions clearly to diverse audiences.
  • Apply medallion architecture, streaming ingestion, and API-based extraction patterns.
  • Write maintainable code and optimize data-flow performance.
  • Contribute to governance frameworks for GxP compliance and data security.

Skills

Data Modeling
ETL/ELT
SQL Proficiency
Streaming Ingestion
Industrial IoT
Cloud (AWS/Azure)
GxP Compliance Knowledge
CI/CD
Data Warehousing
Big Data Concepts

Education

Bachelor's degree in CS/DS/Engineering

Tools

ER/Studio
Erwin
TOAD
DynamoDB
MongoDB
GitHub
Kafka

Job description

Design and build data pipelines for the Lilly Medicine Foundry site
Integrate IT and OT source systems with AWS and Azure cloud data Lakehouse architecture
Capture, ingest, integrate, contextualize, harmonize, and deliver data as reusable data domains and data products
Work across enterprise and edge systems spanning ERP, ELN, MES, LIMS, Historians, PLM, and QMS platforms
Engage with Operations, Quality, Lab, and Engineering stakeholders to understand data needs
Elicit requirements, identify gaps, and translate operational needs into technical designs
Explain architectural decisions to non-technical audiences
Design and develop data pipelines for ingestion, transformation, and integration
Apply medallion architecture, streaming ingestion, and API-based extraction patterns
Write maintainable code and optimize data-flow performance
Participate in design reviews, maintain traceability, and contribute to governance frameworks for GxP compliance
Track emerging AWS and Azure data technologies and contribute to team capability and Tech@Lilly strategy
Analyze complex data domains and develop solutions for analytics
Design, develop, and maintain data solutions for capture, storage, integration, and analytics
Review and recommend data design patterns, performance optimizations, database versions, and deployment strategies

Requirements
  • Bachelor’s degree in Computer Science, Data Science, Engineering, or related field
  • At least 3 years of experience in several disciplines including statistical methods, data modeling, ETL/ELT, ontology development, semantic graph construction and linked data, or relational schema design
  • At least 1 year of experience in a pharmaceutical GxP or scientific environment
  • Experience with AWS and Azure cloud platforms
  • Experience with AI/ML/LLM concepts and tools and building agentic AI solution sets
  • 1–3 years of experience designing large-scale data models for functional, operational, and analytical environments
  • Demonstrated SQL and data modeling proficiency
  • Experience with ER/Studio, Erwin, or TOAD
  • Experience with data streaming, Industrial IoT, MQTT, AMQP, Kafka, and related protocols
  • Understanding of data architecture, data lakehouses, data warehousing, and/or big data concepts
  • Experience with security models and development on large datasets
  • Experience with PostgreSQL, Redshift, Aurora, Athena, Neptune, DynamoDB, MongoDB, and formal database designs
  • Experience with Agile development, CI/CD, GitHub, and automation platforms
  • Knowledge of Data Governance, Master Data Management, and Business Intelligence
  • Prior experience in pharma or another GMP setting
  • Solid knowledge of Computer System Validation processes
  • Problem-solving skills for analyzing, anticipating, and resolving complex issues
  • Learning agility and curiosity
  • Ability to communicate using varied methods in diverse forums
  • Ability to work onsite five days per week in Lebanon, Indiana beginning late 2027–2028; until then, work location is Indianapolis, Indiana
Core Competencies

Demonstrates expertise in designing and building data pipelines, integrating IT and OT systems with AWS and Azure cloud platforms, and applying data governance principles in a pharmaceutical environment. Proficient in data modeling, ETL/ELT processes, and utilizing advanced data technologies for analytics and operational needs.

Highest-signal resume keywords
  • AWS Cloud Platform Experience
  • Azure Cloud Platform Experience
  • Data Pipeline Development
  • SQL Proficiency
  • GxP Compliance Knowledge
ATS Optimization Keywords
Hard Skills
  • Data Modeling
  • ETL/ELT
  • Statistical Methods
  • Data Architecture
  • Data Warehousing
  • AI/ML Concepts
  • Streaming Ingestion
  • Database Design
  • PostgreSQL
  • Agile Development
Soft Skills
  • Problem-Solving Skills
  • Learning Agility
  • Communication Skills
Industry Keywords
  • Pharmaceutical
  • GxP
  • Data Governance
  • Master Data Management
  • Business Intelligence
  • Industrial IoT
  • Big Data Concepts
  • Computer System Validation
Tools & Technologies
  • AWS
  • Azure
  • Kafka
  • GitHub
  • ER/Studio
  • Erwin
  • TOAD
  • DynamoDB
  • MongoDB
  • CI/CD
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff, Data Engineer
Staff, Data Engineer

Jobtailor • Sunnyvale (CA)

On-site
USD 140,000 - 180,000
Principal Scientist, Data Science – R&D, Therapeutics Development & Supply
Principal Scientist, Data Science – R&D, Therapeutics Development & Supply

Jobtailor • Spring House (PA)

On-site
USD 110,000 - 170,000
Associate Director, Data Science, AI Solutions
Associate Director, Data Science, AI Solutions

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Staff Engineer – Data Engineering
Staff Engineer – Data Engineering

Jobtailor • Arizona

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Chicago (IL)

On-site
USD 130,000 - 180,000
Principal Scientist, Data Science – R&D
Principal Scientist, Data Science – R&D

Jobtailor • Spring House (PA)

On-site
USD 90,000 - 130,000
Data Engineer
Data Engineer

Jobtailor • Alabama

On-site
USD 95,000 - 130,000
Data Engineer
Data Engineer

Jobtailor • Maryland

On-site
USD 110,000 - 160,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Oskaloosa (IA)

On-site
USD 110,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Utah

On-site
USD 120,000 - 180,000