Agentic AI Data Engineer - CMC Data Integration

Socket.dev

Indianapolis (IN)

On-site

USD 65,000 - 169,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k)
Pension
Vacation benefits
Medical, dental, vision coverage
Flexible benefits
Life insurance

Job summary

Lilly is seeking an AI Data Engineer to build ingestion pipelines and a unified CMC data backbone. You will implement agent components, mapping, validation, and monitoring to ensure reliable data flows from LIMS/ELN and CDMO partners into a single data backbone.

You will work with scientists and digital architects, owning components end-to-end, and contribute to agentic AI workflows, HITL routing, and validation activities in a regulated environment.

Qualifications

  • MS or BS with required years of hands-on data engineering experience.
  • Proficiency in Python and SQL, with production-quality code skills.
  • Experience building ETL/ELT pipelines from unstructured sources (PDFs, Excel, JSON, XML).
  • Hands-on experience with LLM-powered retrieval-augmented generation and tool-calling.

Responsibilities

  • Build AI-assisted ingestion pipelines for unstructured CDMO/CRO data sources (PDFs, Excel, VO portals).
  • Design data quality framework with automated checks and audit trails.
  • Develop reusable pipeline templates and schema documentation for CDMO partners.

Skills

Python
SQL
ETL/ELT pipelines
LLM-powered apps
Cloud data platforms
Airflow

Education

MS in Computer Science/Engineering or related (1–2 years)
BS in Computer Science/Engineering (3–5 years)

Tools

Azure Data Factory
Databricks
Fabric
S3/Glue/Lambda/Redshift

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Overview

The Bioproduct Research and Development organization strives to deliver creative medicines to patients by developing and commercializing insulins, monoclonal antibodies, novel therapeutic proteins, peptides, oligonucleotide therapies, and gene therapy systems. This multidisciplinary group works collaboratively with our discovery and manufacturing colleagues.

We are seeking an AI Data Engineer to build the data ingestion infrastructure and a unified data model that underpins the modernized CMC Data Backbone. This is a hands-on engineering role with design influence — you will write production-quality pipelines, define CMC data schemas, and work directly with scientists and digital architects to ensure data from internal LIMS/ELN systems and external CDMO partners flows reliably into a single data backbone. You will work with a team of engineers and data scientists. You will have the autonomy to own your components end-to-end. If you want hands-on experience at the intersection of pharmaceutical science and modern agentic AI data engineering — agentic pipelines, document AI, GxP-compliant data infrastructure — this is the role to build that foundation.

Agentic Pipeline Components
  • Implement individual agent components (e.g., document extraction agent, schema mapping agent, validation agent) within the established orchestration framework (LangGraph, LlamaIndex, or equivalent)
  • Write tool-calling logic, handle failure modes, and ensure each agent component is testable and observable with instrumented logging of inputs, outputs, and intermediate decisions
  • Iterate on agent behavior based on real data performance; work with the senior engineer to identify and resolve failure patterns
  • Participate in validation and qualification activities for AI-assisted workflows, supporting documentation that demonstrates computational tools reflect scientific intent
Human-in-the-Loop (HITL) Workflow Implementation
  • Build review queues and flagging logic that surface low-confidence or out-of-specification extractions to scientific reviewers for approval before data is loaded
  • Implement routing logic that captures reviewer decisions, logs outcomes with full audit trail, and reintegrates approved data into the pipeline per 21 CFR Part 11 electronic records requirements
  • Tune flagging thresholds based on feedback from scientific owners; maintain and improve HITL logic as new data sources are onboarded
Data Ingestion & Pipeline Engineering
  • Design and build AI-assisted ingestion pipelines that extract and structure the data from unstructured CDMO/CRO data sources: PDFs (Certificates of Analysis, batch records), Excel files, and vendor portal exports
  • Implement validation, reconciliation, and exception-handling logic to ensure data completeness and integrity before loading
  • Build monitoring and alerting for pipeline health, data quality, and ingestion failures
  • Design a data quality framework with automated checks, rejection handling, and audit trail logging.
  • Develop reusable pipeline templates and schema documentation that reduce onboarding time for new CDMO partners
Required Qualifications
  • MS in Computer Science, Computer Engineering, Data Engineering, or related technical field with 1–2 years of relevant experience; OR
  • BS in Computer Science or Computer Engineering with 3–5 years of hands-on data engineering experience.
  • Proficiency in Python and SQL; ability to write, review, and own production-quality code.
  • Demonstrated experience building ETL/ELT pipelines from unstructured or semi-structured sources (PDFs, Excel, JSON, XML).
  • Hands-on experience building LLM-powered applications: retrieval-augmented generation, tool-calling, multi-step orchestration, or equivalent agentic patterns.
  • Hands-on experience with cloud data platforms: Azure (Data Factory, Databricks, Fabric) or AWS (S3, Glue, Lambda, Redshift).
  • Solid understanding of relational data modeling, schema design, and data normalization principles.
  • Familiarity with data orchestration tools (Airflow, Azure Data Factory, Prefect, or similar)
  • Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1.
Additional Preferences
  • Working knowledge of 21 CFR Part 11, ALCOA+, and GxP data integrity principles, or clear demonstrated ability to apply similar audit/compliance frameworks.
  • Experience integrating data from LIMS, ELN, SDMS, or CDS systems (Benchling, LabVantage, OpenLABS, or equivalent)
  • Familiarity with pharmaceutical CMC data types: analytical results, batch records, stability studies, specifications.
  • Experience with data mesh architecture or data product ownership models.
  • Knowledge of MLOps practices and preparing data for AI/ML model training in regulated environments.
  • Exposure to regulatory submission data formats (eCTD, CTD, CDISC SEND/SDTM).
  • Experience with CI/CD pipelines (GitHub Actions, Azure DevOps) applied to data engineering workloads.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form https://careers.lilly.com/us/en/workplace-accommodation for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.

Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is

$65,250 - $169,400

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance).

  • eligibility to participate in a company-sponsored 401(k)
  • pension
  • vacation benefits
  • eligibility for medical, dental, vision and prescription drug benefits
  • flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts)
  • life insurance and death benefits
  • certain time off and leave of absence benefits
  • well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities)

Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Data Engineer - CMC Data Integration
Agentic AI Data Engineer - CMC Data Integration

BioSpace • Indianapolis (IN)

On-site
USD 65,000 - 169,000
401(k) plan
Pension
Vacation
Senior Advisor, Agentic AI Solutions and SciML Engineer
Senior Advisor, Agentic AI Solutions and SciML Engineer

BioSpace • Indianapolis (IN)

On-site
USD 129,000 - 209,000
Company bonus
Health, dental, vision benefits
401(k) plan
Senior Advisor, Agentic AI Solutions and SciML Engineer
Senior Advisor, Agentic AI Solutions and SciML Engineer

Initial Therapeutics, Inc. • Indianapolis (IN)

On-site
USD 129,000 - 209,000
401(k) plan
Pension
Vacation
+1
Senior Advisor, Agentic AI Solutions and SciML Engineer
Senior Advisor, Agentic AI Solutions and SciML Engineer

Socket.dev • Indianapolis (IN)

On-site
USD 129,000 - 209,000
401(k) plan
Pension
Medical, dental, vision
+2
Advisor, Data Scientist - CMC Data Products
Advisor, Data Scientist - CMC Data Products

Eli Lilly and Company • Indianapolis (IN)

On-site
USD 126,000 - 245,000
Comprehensive benefit program
401(k) and pension options
Flexible benefits and wellness programs
Advisor, Data Scientist - CMC Data Products
Advisor, Data Scientist - CMC Data Products

BioSpace • Indianapolis (IN)

On-site
USD 126,000 - 244,000
401(k) matching
Pension
Medical insurance
+3
Sr. Principal or Engineering Advisor - Agentic Lab Automation Integration
Sr. Principal or Engineering Advisor - Agentic Lab Automation Integration

Eli Lilly and Company • San Diego (CA)

On-site
USD 132,000 - 223,000
Comprehensive benefit program
Company-sponsored 401(k)
Flexible benefits
Machine Learning & Data Operations Engineer
Machine Learning & Data Operations Engineer

Eli Lilly and Company • Indianapolis (IN)

On-site
USD 152,000 - 244,000
401(k)
Pension
Vacation benefits
+5
Data Architect, Data Foundry
Data Architect, Data Foundry

BioSpace • San Francisco (CA)

On-site
USD 132,000 - 194,000
401(k) plan
Pension
Life insurance
+1
Technical Lead - Software Developer, Data Foundry
Technical Lead - Software Developer, Data Foundry

BioSpace • San Francisco (CA)

On-site
USD 152,000 - 244,000
401(k) plan
Medical, dental, vision benefits
Pension
+4