Data Engineer

100 Eli Lilly and Company

South San Francisco (CA)

Hybrid

USD 158,000 - 231,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

401(k)
Pension
Vacation benefits
Medical, dental, vision
Flexible benefits
Life insurance
Time off benefits

Job summary

Eli Lilly and Company in Silicon Valley seeks a Data Engineer to build and maintain data platforms powering AI-driven research. You will ingest, transform, and deliver chemical, biological, and experimental data for machine learning workflows, partnering with AI scientists and lab researchers to ensure data quality and accessibility.

The role requires strong Python, SQL, and data modeling skills, plus experience with Spark, Airflow, and cloud environments.

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Engineering, Mathematics, or a related technical field.
  • 5+ years of data engineering experience building and operating production data systems.
  • Experience building scalable data platforms for AI/ML workflows.

Responsibilities

  • Build scalable data pipelines for AI-driven research and discovery.
  • Ensure data accuracy, traceability, and accessibility at scale.
  • Collaborate with AI scientists, engineers, and lab researchers to enable data-driven experiments.

Skills

Python
SQL
Data modeling
Distributed data processing
CI/CD
Git
Big data

Education

Bachelor's degree in CS/Math/Engineering/Data Science

Tools

Airflow
Dagster
Spark
Ray
Dask
Kafka

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life‑changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley!

Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us!

About the Lilly and NVIDIA Partnership

Lilly and NVIDIA are launching a new AI co‑innovation lab in the heart of Silicon Valley — an up‑to‑$1billion, multi‑year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose‑built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.

What You’ll Be Doing

As a Data Engineer, you will build and maintain the data platforms that power AI‑driven research and discovery. You will develop scalable pipelines that ingest, transform, and deliver chemical, biological, and experimental data for machine learning and scientific workflows. Partnering with AI Scientists, AI Engineers, and laboratory researchers, you will ensure that data is accurate, traceable, and accessible at scale. Your work will provide the trusted data foundation behind next‑generation AI models and experiments.

How You’ll Succeed

Engineer datasets in large language environment for model training specifically efficient formats and storage layout (Parquet, Zarr, Arrow) and delivery fast enough that GPU clusters are never left waiting on data. Design, develop, and maintain scalable and efficient data pipelines to support data analytics, reporting, and machine learning initiatives. Ensure seamless data flow between systems and applications, optimizing data transfer and transformation processes for performance and scalability. Build the ingestion path from the automated lab, so experimental results reach the models in hours rather than weeks, closing the loop between what a model proposes and what the next model learns from. Own the correctness of what models train on completeness, sound joins across experimental sources, and validation that catches a bad dataset before it reaches a training run rather than after. Build dataset versioning, lineage, and reproducibility into the platform, so any model can be traced to the exact data it was trained on months or years later. Work with the laboratory, instrument, and external teams producing the data so that a change upstream does not quietly corrupt a training run downstream.

What You Should Bring

Strong Python, or equivalent experience building data‑intensive software systems. Strong SQL and data modeling experience including designing schemas that hold up as scientific data grows and diversifies, with expert knowledge of Postgres or a comparable enterprise database. Distributed data processing (Spark, Ray, or Dask) and pipeline orchestration (Airflow or Dagster) at scale. Experience with cloud platforms — AWS and Azure preferred — and with high‑performance and object storage feeding large‑scale compute environments. A track record of building data systems that other people depend on, and of taking responsibility for them when they broke. Strong testing practices and test automation, with solid CI/CD and Git fundamentals. Adaptability and a collaborative mindset, with the ability to translate complex scientific questions into data solutions that accelerate experimentation and decision‑making. Experience streaming and event‑driven integration (Kafka, MQTT, or AMQP), including instrument and laboratory data capture. Cheminformatics or scientific data experience — compound registration, structure notation (SMILES, InChI, HELM), RDKit, multi‑omics, assay, or sequencing data — is a strong plus. Prior experience across the following: data modeling, ETL/ELT at scale, ontology development, semantic graph construction and linked data, or relational schema design. Experience standing up, migrating, or consolidating databases and data platforms, including production cutover of systems in active use.

Your Basic Qualifications

Bachelor’s degree in Computer Science, Data Science, Engineering, Mathematics, or a related technical field. 5+ years of data engineering experience building and operating production data systems.

Location & Work Flexibility

This role is based at our Silicon Valley Hub. We offer a flexible hybrid work model, with three days onsite and two days working remotely each week, supporting both collaboration and work‑life balance.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form https://careers.lilly.com/us/en/workplace-accommodation for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.

Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees.

  • Africa
  • Middle East
  • Central Asia (AMECA)
  • Black Employees at Lilly (BE@Lilly)
  • Chinese Culture Network (CCN)
  • EnAble
  • Evolve
  • Lilly Indian Network (LIN)
  • Organization of Latinx at Lilly (OLA)
  • Pride (LGBTQ+ Allies)
  • Veterans Leadership Network (VLN)
  • Women’s Initiative for Leading at Lilly (WILL)

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is $157,500 - $231,000. Full‑time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance).

In addition, Lilly offers a comprehensive benefit program to eligible employees,

  • including eligibility to participate in a company-sponsored 401(k)
  • pension
  • vacation benefits
  • eligibility for medical, dental, vision and prescription drug benefits
  • flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts)
  • life insurance and death benefits
  • certain time off and leave of absence benefits
  • well‑being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities)

We hope that you seek to join us on our journey as we create medicine and deliver improved outcomes for patients across the globe!

#WeAreLilly At Lilly we strive to ensure our employees are part of a team that cares about them and our shared purpose of making life better for those around the world. How do we do this? We continue to look for ways to include, innovate, accelerate and deliver while maintaining integrity, excellence and respect for people.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Scientist (Model Building & Training)
AI Scientist (Model Building & Training)

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 168,000 - 268,000
401(k) match
Pension plan
Vacation benefits
+5
AI Engineer
AI Engineer

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 253,000
401(k)
Pension
Vacation benefits
+1
Sr. AI Science Lead
Sr. AI Science Lead

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 260,000 - 381,000
Company bonus
401(k) and benefits
Health, dental and vision
HPC Systems Administrator
HPC Systems Administrator

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+1
AI Scientist (Model Building & Training)
AI Scientist (Model Building & Training)

Initial Therapeutics, Inc. • San Francisco (CA)

Hybrid
USD 168,000 - 268,000
Company bonus
401(k) plan
Health, dental, vision benefits
+4
Technical Lead - Software Developer, Data Foundry
Technical Lead - Software Developer, Data Foundry

BioSpace • San Francisco (CA)

On-site
USD 152,000 - 244,000
401(k) plan
Medical, dental, vision benefits
Pension
+4
Data Engineer
Data Engineer

BioSpace • Indianapolis (IN)

On-site
USD 65,000 - 158,000
Bonus program
Comprehensive benefits
401(k)
Sr. AI Science Lead
Sr. AI Science Lead

Eli Lilly and Company • San Francisco (CA)

Hybrid
USD 260,000 - 381,000
Company bonus
Health benefits (medical, dental, vis.
401(k) and pension
+2
Postdoctoral Fellow - Agentic AI Solutions and SciML
Postdoctoral Fellow - Agentic AI Solutions and SciML

100 Eli Lilly and Company • United States

On-site
USD 58,000 - 123,000
401(k) plan
Pension
Vacation benefits
+1
Agentic AI Data Engineer - CMC Data Integration
Agentic AI Data Engineer - CMC Data Integration

BioSpace • Indianapolis (IN)

On-site
USD 65,000 - 169,000
401(k) plan
Pension
Vacation