Senior Platform Data Engineer

Geisinger

Danville (PA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A healthcare organization in Danville, Pennsylvania, is seeking a Senior Platform Data Engineer to own the data infrastructure for AI initiatives. You'll stream and curate clinical data in Databricks, design ingestion pipelines, and manage embedding and retrieval pipelines. A minimum of 5 years in data engineering and expert-level knowledge of Databricks is required. Candidates with experience in real-time ingestion frameworks and familiarity with clinical data sources will stand out in this role.

Qualifications

  • 5+ years in data engineering with experience in batch and streaming data pipelines.
  • Expert-level skills in Databricks: Delta Live Tables, PySpark, Unity Catalog, Feature Store.
  • Hands-on experience with real-time data ingestion frameworks.

Responsibilities

  • Streams clinical data into Databricks for model use.
  • Curates shared clinical feature tables.
  • Designs document ingestion pipelines and manages embedding pipelines.

Skills

Data pipeline development
Databricks (Delta Live Tables, PySpark)
Real-time data ingestion
SQL
Python (pandas)
Data governance principles

Education

Bachelor's Degree in Related Field
Master's Degree in Related Field

Tools

Vector databases (Pinecone, Weaviate, etc.)
Epic SDE / epic-ws
CDIS Data Warehouse

Job description

Job Summary

The Senior Platform Data Engineer owns roadmap, priorities, platform standards, and architecture reviews; provides formal input on performance reviews. This position makes clinical data ready for AI at scale: owning the shared data products, retrieval infrastructure, and platform administration that the entire AI portfolio depends on. Owns real‑time data feeds, reusable clinical data models and feature pipelines, RAG retrieval infrastructure, and Databricks platform administration.

Job Duties
  • Streams data from Epic SDE, ADT feeds, lab results, and other clinical sources into Databricks for downstream model consumption.
  • Curates shared clinical feature tables (patient demographics, labs, vitals, diagnoses, utilization history, imaging metadata) in Databricks/Unity Catalog that multiple AI programs consume for model training, validation, and monitoring.
  • Owns RAG infrastructure, the shared retrieval‑augmented generation platform that agentic and generative AI programs use to ground LLM outputs in organizational knowledge.
  • Designs and operates document ingestion pipelines: normalizing clinical documents, policies, guidelines, and unstructured data sources into formats ready for embedding and retrieval.
  • Implements and optimizes chunking strategies tailored to healthcare content (e.g., preserving clinical note structure, section‑aware chunking for guidelines and protocols).
  • Manages the embedding pipeline: selecting, tuning, and versioning embedding models (domain‑specific clinical models where they outperform general‑purpose).
  • Administers the vector database: schema design, indexing, metadata management, access controls, and performance tuning.
  • Builds and maintains retrieval pipelines: hybrid search (vector + keyword/BM25), reranking, and relevance filtering to maximize retrieval precision for downstream agents and LLM applications.
  • Establishes data quality gates for RAG: automated profiling, completeness checks, and accuracy scoring before content enters the vector store.
  • Monitors retrieval quality metrics (Precision@K, Recall@K, MRR) and continuously optimizes retrieval performance.
  • Databricks workspace configuration and Unity Catalog governance.
  • Cluster policies, compute management, and cost monitoring.
  • Manges user/group management and access control.
  • Administrator for Feature Store.
Key Technologies
  • Databricks (Delta Live Tables, Feature Store, PySpark, Unity Catalog)
  • Epic SDE / epic-ws for real‑time clinical data extraction
  • Vector databases (Pinecone, Weaviate, Qdrant, or Databricks Vector Search)
  • Embedding models and pipelines (clinical domain‑specific and general‑purpose)
  • SQL, pandas
  • Streaming and batch ingestion patterns
  • CDIS Data Warehouse (source system for batch clinical data)
Required Skills & Qualifications
  • 5+ years in data engineering, with strong experience building both batch and streaming data pipelines
  • Expert‑level Databricks skills: Delta Live Tables, PySpark, Unity Catalog, Feature Store
  • Hands‑on experience with real‑time data ingestion (Kafka, Spark Structured Streaming, or comparable frameworks)
  • Strong SQL and Python (pandas, PySpark) skills for data transformation and feature engineering
  • Experience administering Databricks workspaces: cluster policies, compute management, access controls, cost monitoring
  • Familiarity with clinical data models and healthcare data sources (EHR extracts, ADT feeds, lab results, claims data) strongly preferred
  • Experience with Epic data extraction methods (SDE, FHIR, epic‑ws) a significant plus
  • Understanding of data governance principles: lineage, quality monitoring, access controls
Education

Bachelor's Degree‑Related Field of Study (Required), Master's Degree‑Related Field of Study (Preferred)

Experience

Minimum of 5 years‑Relevant experience (Required)

We are proud to be an affirmative action, equal opportunity employer and all qualified applicants will receive consideration for employment regardless to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or status as a protected veteran.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Data Engineer: Clinical AI Data Infra
Senior Platform Data Engineer: Clinical AI Data Infra

Geisinger • Danville (PA)

On-site
USD 100,000 - 130,000
Senior Manager, Data Engineer, Clinical Operations
Senior Manager, Data Engineer, Clinical Operations

Scorpion Therapeutics • Princeton (NJ)

On-site
USD 154,000 - 186,000
Health Coverage
Wellbeing support
401(k)
Lead Software Engineer (Databricks Platform Engineer)
Lead Software Engineer (Databricks Platform Engineer)

Vizient, Inc • Irving (TX)

On-site
USD 117,600 - 206,000
Data Engineer
Data Engineer

Scorpion Therapeutics • Indianapolis (IN)

On-site
USD 120,000 - 180,000
Data Engineer (AI/ML)
Data Engineer (AI/ML)

001_BCBSA Blue Cross and Blue Shield Association • Chicago (IL)

On-site
USD 100,000 - 139,000
Paid time off
Medical/dental/vision insurance
Generous 401(k) matching
+1
Principal Data Engineer
Principal Data Engineer

codametrix • United States

On-site
USD 140,000 - 210,000
EPIC Data Platform Lead
EPIC Data Platform Lead

Nityo Infotech • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Sr. Data Engineer (Databricks)
Sr. Data Engineer (Databricks)

TurningPoint Healthcare Solutions • Town of Florida (NY)

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Analytica • Washington

On-site
USD 140,000 - 180,000
Bonuses
Employer-paid health care
Training funds
+1
Senior Data Engineer
Senior Data Engineer

analyticallc • Washington

On-site
USD 130,000 - 190,000
Bonuses
Employer-paid health care
Training & development funds
+1