Pharma Data Engineer - Databricks AWS

RADcube

Indianapolis (IN)

Hybrid

USD 110,000 - 150,000

Full time

32 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

RADcube is seeking a Pharma Data Engineer to join our pharma-focused data team in a hybrid role in Indianapolis. The candidate will design, build, and govern data infrastructure powering analytics and reporting, with significant business-facing responsibilities.

The role combines hands-on pipeline work with executive-level stakeholder interaction, data governance, and green-field data domain development in a regulated industry, using Databricks and AWS.

Qualifications

  • 3–5 years of data engineering experience in pharma industry.
  • ETL/ELT development for batch and streaming pipelines.
  • Advanced SQL proficiency.
  • Python or Scala for data transformation.
  • API integration for external data ingestion.
  • Experience with data pipeline orchestration (Airflow, Databricks Workflows, AWS Glue).
  • Familiarity with GxP-regulated data environments and data privacy compliance (21 CFR Part 11, GDPR).
  • Experience with Databricks and AWS in pharma context.
  • Ability to translate business needs into technical specs and present to executives.

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines (batch and streaming).
  • Write and optimize advanced SQL and develop data transformations in Python/Scala.
  • Integrate external data sources via APIs and manage pipeline orchestration (Airflow, Databricks Workflows, AWS Glue).
  • Apply data quality, governance, and lineage practices in regulated environments.
  • Collaborate with stakeholders across R&D, Manufacturing, Quality, and Commercial to translate requirements into specs.
  • Present technical work and data strategy to executive audiences.
  • Help stand up new data domains from scratch (green-field builds) and drive adoption of new solutions.

Skills

ETL/ELT development
Advanced SQL
Python or Scala
Data orchestration
APIs integration
Databricks
AWS
Tableau/Power BI

Tools

Databricks (Delta Lake)
AWS (S3, Glue)
Airflow
Databricks Workflows

Job description

Hybrid – Indianapolis, IN
We are seeking a

Pharma Data Engineer – Databricks & AWS
Hybrid – Indianapolis, IN
We are seeking a Data Engineer with 3–5 years of experience working specifically within the pharma industry to join a pharma-focused data team. This is a senior-flavored engineering role that combines hands‑on pipeline and platform work with significant business‑facing responsibility — including translating business needs into technical specs, presenting to executive‑level stakeholders, and helping stand up new data domains from the ground up. You will design, build, and govern the data infrastructure that powers analytics and reporting across the business, while also acting as a trusted technical partner to non‑technical stakeholders.

Key Responsibilities

Design, build, and maintain scalable ETL/ELT pipelines (batch and streaming) using Databricks, AWS, and related orchestration tools. Write and optimize advanced SQL, and build data transformations in Python or Scala. Integrate external data sources via APIs and manage pipeline orchestration (Airflow, Databricks Workflows, AWS Glue). Apply data quality, governance, cataloging, and lineage practices aligned with regulated‑industry standards. Work within GxP‑regulated data environments and apply awareness of data privacy/compliance considerations (e.g., 21 CFR Part 11, GDPR where applicable). Partner with business stakeholders across the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development) to gather and translate requirements into technical specifications. Present technical work and data strategy to executive‑level audiences. Prioritize high‑impact data initiatives and proactively identify and avoid duplicated data efforts. Support change management and adoption of new data solutions across business teams. Help stand up new data domains from scratch (green‑field build), not just maintain existing ones.

Requirements
Required Qualifications
Data Engineering & Pipelines
  • Data Engineering & Pipelines
  • ETL/ELT development (batch and streaming)
  • Advanced SQL (joins, window functions, query optimization)
  • Python or Scala for data transformation
  • Data pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)
  • API integration for external data source ingestion
Platforms & Tools
  • Databricks (Delta Lake, Unity Catalog, Genie)
  • Cloud platforms — AWS (S3, Glue, Athena) and/or Azure/Google Cloud Platform equivalents
  • Data warehousing concepts (dimensional modeling, star schema)
  • BI/visualization tools (Tableau, Power BI, or similar) to understand downstream consumption
Data Quality & Governance
  • Data profiling and cleansing techniques
  • Metadata management and data cataloging
  • Master data management (MDM) principles
  • Data lineage tracking
  • Data governance frameworks (especially regulated‑industry standards)
Pharma / Life Sciences Domain Knowledge
  • Familiarity with GxP‑regulated data environments
  • Understanding of the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development)
  • Awareness of data privacy/compliance considerations (21 CFR Part 11, GDPR where applicable)
  • Knowledge of common pharma data domains (clinical, manufacturing, quality, commercial)
Stakeholder Management
  • Requirements gathering and translation (business need technical spec)
  • Cross-functional communication (Business IT)
  • Executive‑level presentation skills (given EC visibility)
  • Change management / adoption support
Analytical & Strategic Thinking
  • Prioritization frameworks (identifying high‑impact vs. low‑value data asks)
  • Cost‑avoidance mindset (spotting duplication before it happens)
  • Ability to work with ambiguity and evolving priorities
Project & Program Skills
  • Agile/Scrum familiarity
  • Documentation discipline (data dictionaries, source‑to‑target mappings)
  • Vendor/partner coordination (if external data sources are involved)
Nice‑to‑Have Differentiators
  • Prior consulting or client‑facing delivery experience
  • Experience standing up new data domains from scratch (green‑field vs. maintenance)
  • Familiarity with AI/GenAI‑enabled analytics tools
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Jobtailor • United States

On-site
USD 120,000 - 165,000
Data Engineer
Data Engineer

Scorpion Therapeutics • Indianapolis (IN)

On-site
USD 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Tiger Analytics • Chicago (IL)

On-site
USD 130,000 - 180,000
Innovation, Data & Analytics Team, Data Engineer
Innovation, Data & Analytics Team, Data Engineer

Scorpion Therapeutics • Pennsylvania

Hybrid
USD 100,000 - 180,000
Technical Life Sciences Consultant – Databricks & Commercial Pharma
Technical Life Sciences Consultant – Databricks & Commercial Pharma

Bestica Inc. • Princeton (NJ)

On-site
USD 140,000 - 190,000
Lead Data Engineer
Lead Data Engineer

MathCo • New Jersey

On-site
USD 100,000 - 140,000
Data Engineer (P-175)
Data Engineer (P-175)

Smash CR • Dallas (TX)

On-site
USD 110,000 - 140,000
Digital Platforms, Senior Data Engineer - PCI Pharma Services
Digital Platforms, Senior Data Engineer - PCI Pharma Services

OpenTalent • Philadelphia

On-site
USD 130,000 - 170,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Naperville (IL)

On-site
USD 120,000 - 160,000