Data Engineer

Scorpion Therapeutics

Indianapolis (IN)

On-site

USD 120,000 - 180,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Scorpion Therapeutics is seeking a data engineer to design, build, and optimize scalable data pipelines across the CE data ecosystem. You will work with Databricks, PySpark, Python, SQL, and Delta Lake to ingest, transform, and store data in data warehouses and data lakes.

You will implement the Bronze–Silver–Gold medallion architecture, establish data governance, and enable secure access to PHI while collaborating with analytics and product teams in an Agile environment.

Qualifications

  • Bachelor’s degree in CS/Engineering/IS or related quantitative field.
  • 5+ years data engineering/ETL development.
  • Proficient in SQL and at least one language (Python/Java/Databricks).
  • Experience with cloud data platforms (Databricks, AWS/Azure/GCP) and services (S3, Redshift, Snowflake, ADLS, BigQuery).
  • Proficient with Git-based CI/CD workflows (GitHub Actions or equivalent).

Responsibilities

  • Design, develop, and optimize scalable data pipelines using Databricks and Python.
  • Implement canonical data models across medallion architecture Bronze → Silver → Gold.
  • Evaluate/apply Databricks capabilities for performance, cost, and scalability.
  • Implement and maintain ELT/ETL workflows (Databricks Workflows, Auto Loader, Structured Streaming, Delta Live Tables).
  • Build CI/CD pipelines for CE data and artifacts.
  • Automate data ingestion and product creation to reduce manual maintenance and onboarding.

Skills

Data engineering
SQL
Python / Databricks
Cloud platforms

Education

Bachelor’s degree in CS/Engineering/IS or related field

Tools

Databricks
AWS/Azure/GCP
S3 / Redshift
Snowflake / BigQuery

Job description

What You Will Do
Data Engineering & Pipeline Development
  • Design, develop, and optimize scalable data pipelines using Databricks, PySpark, Python, SQL, and Delta Lake to ingest, transform, and load data into data warehouses and data lakes.
  • Build Databricks pipelines implementing canonical data models across the medallion architecture (Bronze → Silver → Gold) within the CE trust boundary.
  • Evaluate/apply Databricks capabilities (Unity Catalog, Delta Lake, Databricks Workflows, serverless compute, Lakebase, ingestion connectors) based on performance, cost, and scalability.
  • Implement and maintain ELT/ETL workflows (Databricks Workflows, Auto Loader, Structured Streaming, Delta Live Tables).
  • Build and maintain CI/CD pipelines (GitHub Actions; dev → test → prod promotion) for CE data and artifacts.
  • Automate data ingestion and product creation to reduce manual maintenance and onboarding.
Data Governance, Quality & Security
  • Implement data governance policies ensuring data quality, integrity, security, and compliance (e.g., GxP, HIPAA) and covered-entity constructs.
  • Implement row/column-level security, masking, and tokenization to enforce PHI isolation.
  • Use Unity Catalog for metadata, lineage, and access control.
  • Establish testing/validation (pytest, DLT/Great Expectations) and monitoring/alerting.
Data Modeling & Architecture
  • Develop/maintain data models, schemas, metadata; follow Lakehouse/Medallion principles.
  • Create reusable transformation frameworks and automated data quality checks.
  • Partner on reference architecture; document pipelines/processes.
Collaboration & Innovation
  • Translate stakeholder data requirements into technical solutions.
  • Monitor performance, troubleshoot, and improve availability/reliability.
  • Participate in code reviews/architecture discussions; promote best practices.
  • Evaluate/recommend new tools; support production integration of ML/AI tools with PHI classification/consent.
  • Participate in Agile ceremonies (Jira or equivalent).
Your Minimum Basic Qualifications
  • Bachelor’s degree in CS/Engineering/IS or related quantitative field.
  • 5+ years data engineering/ETL development.
  • Proficient in SQL and at least one language (e.g., Python/Java/Databricks).
  • Experience with cloud data platforms (Databricks, AWS/Azure/GCP) and services (S3, Redshift, Snowflake, ADLS, BigQuery).
  • Proficient with Git-based CI/CD workflows (GitHub Actions or equivalent).
What You Should Bring
  • Excellent problem-solving; ability to build/test pipelines from architecture.
  • Strong communication/collaboration.
  • Prior pharma/life sciences experience (preferred).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Recru, LLC. • Sugar Land (TX)

On-site
USD 120,000 - 150,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Databricks Engineer
Databricks Engineer

CMT Services, Inc. • Adelphi (MD)

On-site
USD 100,000 - 130,000
Data Solutions Engineer
Data Solutions Engineer

Jobtailor • Durham (NC)

On-site
USD 110,000 - 160,000
BI Data Engineer II
BI Data Engineer II

Jobtailor • Boston (MA)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

Recru, LLC. • Spring (TX)

On-site
USD 180,000 - 230,000
Data Architect
Data Architect

Pho Prime, LLC • Shelton (CT)

On-site
USD 120,000 - 190,000
Mobility Allowance
Databricks Engineer
Databricks Engineer

iLink Digital • Milpitas (CA), Northern (KY)

On-site
USD 120,000 - 160,000
Sr Data Engineer
Sr Data Engineer

Hertz • Estero (FL), Northern (KY)

Hybrid
USD 117,000 - 143,000
Data Engineer
Data Engineer

Ranger Technical Resources • Town of Florida (NY)

On-site
USD 130,000 - 185,000