Senior Data Architect AWS And Databricks Modernization

Eli Lilly and Company

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Eli Lilly and Company in Bengaluru seeks a Senior Data Architect to lead Databricks lakehouse modernization at scale, combining hands-on implementation with architectural leadership to migrate legacy data warehouses to a unified Databricks platform for Clinical and Non-Clinical data.

You will design reusable data products, own Unity Catalog governance, and optimize performance and cost, collaborating across India to deliver scalable, auditable data solutions that accelerate drug discovery and

Qualifications

  • 5+ years hands-on experience migrating or modernizing legacy data warehouses/on-prem platforms.
  • Hands-on Databricks lakehouse modernization experience preferred.

Responsibilities

  • Migrate legacy data warehouses to Databricks lakehouse at scale.
  • Build migration pipelines and accelerators for schema conversion and backfills.
  • Define reusable modernization patterns and governance for enterprise-scale use.
  • Architect and optimize the Databricks lakehouse across multiple domains.
  • Lead AI-assisted tooling, data quality, and analytics-ready data initiatives.

Skills

Databricks
Lakehouse
Delta Live Tables
PySpark
Unity Catalog
AWS
Data governance

Job description

Job Description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

CTI MD Tech@Lilly
Senior Data Architect — AWS & Databricks Modernization at Scale
Position Description
About Lilly

At Lilly, everything we do starts with patients. We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it. If you're driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly.

About Technology@Lilly

At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers. Lilly in Bengaluru builds the capabilities that make this possible, cloud platforms, AI systems, and automation at enterprise scale, all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve.

About The Business Function

The Clinical & Non-Clinical Data Organization at Eli Lilly and Company is responsible for the design, build, and operation of enterprise data platforms that power drug discovery, clinical development, and regulatory submissions. Data Hub is building a robust Data Strategy to make Lilly's Clinical and Non-Clinical data AI-ready and audit-ready, delivering scalable, governed, and reusable data products that accelerate how medicines reach patients. Our data engineering organization sits at the intersection of science, technology, and patient impact — connecting Clinical and Non-Clinical data across the full chain, from ingestion to consumption.

Role
Senior Data Architect — AWS + Databricks Modernization at Scale

The Senior Data Architect (R5) is a hands-on Databricks and lakehouse modernization leader who owns the architecture and execution of large-scale migrations from legacy warehouses and point platforms onto a unified Databricks lakehouse for the Clinical and Non-Clinical data domain. This is a builder role: the architect writes code, builds migration pipelines, and personally ships production lakehouse assets — in addition to defining the modernization roadmap and rollout patterns that scale across dozens of domains. The role also owns the roadmap for semantic modeling and analytics-ready data, builds reusable data products the organization can adopt, drives data quality on clinical data specifically, and defines the metrics that enable data-driven decision making. AI-assisted and agentic ways of working are the expected default. Domain knowledge of the clinical data landscape is an added advantage. This role is split 70% hands-on technical execution including coding and 30% strategy, and shapes the future technology landscape.

Key Responsibilities
Lakehouse Modernization & Migration at Scale (Hands-On)
  • Own the technical roadmap and execution for migrating legacy warehouses, on-prem databases, and point data platforms onto Databricks — sequencing dozens of Clinical and Non-Clinical domains with minimal disruption.
  • Personally build migration pipelines and re-platforming accelerators (schema conversion, historical backfill, dual-run validation, cutover automation) that move data at scale with parity and auditability.
  • Define reusable modernization patterns — landing zone design, medallion (bronze/silver/gold) conventions, workspace/catalog topology — that scale consistently as new domains onboard.
  • Right-size and standardize the Databricks platform footprint across environments (dev/test/prod, multiple workspaces) for cost, performance, and governance at enterprise scale.
  • Establish cutover, rollback, and data-reconciliation practices that de-risk large-scale migrations in a regulated environment.
Databricks Lakehouse Architecture & Engineering (Hands-On)
  • Architect and build the Clinical/Non-Clinical lakehouse on Databricks, applying medallion design across Delta Lake tables at scale across multiple domains.
  • Personally build and optimize Delta Live Tables (DLT) pipelines, Databricks Workflows, and PySpark/Spark SQL jobs for high-volume ingestion, transformation, and curation.
  • Own Unity Catalog design and rollout — catalogs, schemas, access control, lineage, and data sharing — as the governance backbone across an expanding domain footprint.
  • Tune performance and cost at scale: Photon, cluster policies, job/task orchestration, auto-scaling, and Databricks SQL Serverless warehouses across many concurrent workloads.
  • Package and deploy pipelines using Databricks Asset Bundles and CI/CD; evaluate Lakehouse Federation and cross-workspace patterns for enterprise-wide access.
AI-Native & Agentic Data Engineering
  • Use AI-assisted and agentic tooling by default — migration gap analysis, pipeline scaffolding, code conversion, and documentation — to accelerate modernization at scale.
  • Apply Databricks Mosaic AI / MLflow and LLM-based tooling to automate schema mapping, ontology alignment, and data-quality scoring during migration.
  • Build reusable AI-assisted accelerators that other engineers and architects reuse to modernize additional domains faster.
Semantic Modeling, Analytics-Ready Data & Reusable Data Products
  • Define and own the multi-quarter roadmap for semantic modeling and analytics-ready data — ontologies, taxonomies, dimensional models, and semantic layers — across Clinical and Non-Clinical domains.
  • Design and build reusable, governed data products (documented, discoverable, versioned) on the lakehouse that other squads can adopt directly rather than rebuilding equivalents.
  • Own data quality specifically for clinical data — define quality rules, thresholds, and remediation workflows for clinical data sets, and track quality trends over time.
  • Define the metrics and KPIs (data quality scores, product adoption/reuse rates, pipeline reliability, time‑to‑insight) that enable data-driven decision making across the India data organization, and report progress against the roadmap.
Governance, Standards & Automation
  • Implement role-based and attribute-based access control and encryption for data at rest and in transit within Unity Catalog, consistently across every migrated domain.
  • Stand up and maintain the data standards platform — the reference implementation, templates, and lineage/quality practices for how Clinical and Non-Clinical data is modeled and migrated.
  • Automate lineage capture, schema-drift detection, and post‑migration data-quality validation so the team scales modernization without proportional manual effort.
Collaboration & Stakeholder Engagement
  • Partner with business SMEs, solution architects, and engineering teams to sequence and translate legacy platform needs into Databricks‑based modernization plans.
  • Communicate migration risk, trade‑offs, and rollout progress clearly to technical and non‑technical stakeholders, and own the India lakehouse modernization roadmap.
Qualifications Required
Required — Lakehouse Modernization at Scale (Must‑Have, Hands‑On)
  • 5+ years hands‑on experience migrating or modernizing legacy data warehouses/on‑prem platforms (
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Architect – AWS & Databricks Modernization
Senior Data Architect – AWS & Databricks Modernization

Eli Lilly and Company • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Senior Data Architect – AWS & Databricks Modernization
Senior Data Architect – AWS & Databricks Modernization

550 Eli Lilly Services India Pvt Ltd • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Resident Solution Architect
Resident Solution Architect

Celebal Technologies • Dadri

On-site
INR 3,000,000 - 6,000,000
Walk-in | Data Architect
Walk-in | Data Architect

Celebal Technologies • Pune District

On-site
INR 4,000,000 - 8,000,000
Lead Data Engineer
Lead Data Engineer

Celebal Technologies • Dadri, Jaipur, Bengaluru

On-site
INR 3,000,000 - 7,200,000
Solutions Architect
Solutions Architect

Celebal Technologies • Bengaluru

On-site
INR 2,000,000 - 2,500,000
Principal Architect - Databricks
Principal Architect - Databricks

Unisys • Bengaluru

On-site
INR 5,000,000 - 7,000,000
Director, Enterprise Data & Analytics Architect - Lilly USA Commercial Technology
Director, Enterprise Data & Analytics Architect - Lilly USA Commercial Technology

550 Eli Lilly Services India Pvt Ltd • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Lead Technical Architect – Data Platform Modernization
Lead Technical Architect – Data Platform Modernization

Tata Consultancy Services • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Databricks Developer / Data Architect
Databricks Developer / Data Architect

ExlService Holdings, Inc. • India

Hybrid
INR 1,500,000 - 2,100,000