Senior Translational Data and AI Engineer

Creative Solutions Services, LLC

Wilmington (DE)

On-site

USD 140,000 - 180,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kaztronix is seeking a contract Senior Translational Data and AI Engineer to modernize ingestion and delivery of biomarker and clinical data on a centralized AWS-native lakehouse. You will build end-to-end pipelines, tests, and guardrails while collaborating with the lead engineer on a fast, reproducible stack.

The role emphasizes translating scientific requirements into durable data models and scalable workflows, including legacy R/PySpark integration and AI-assisted tooling.

Qualifications

  • AI-native engineering experience building AI-assisted pipelines and guardrails.
  • Bachelor's or master's degree in CS, Data Eng, Bioinformatics, or related field.
  • 4+ years of data engineering with production pipelines on AWS (S3, ECS/Fargate, Redshift or equivalent).
  • Strong Python and SQL with modern data engineering libraries.
  • Hands-on with dbt and workflow tools (Dagster, Airflow, or Prefect).
  • Data quality instincts to catch silent failures and ensure complete deliveries.
  • Familiarity with lakehouse architectures, ETL, and multi-modal data schemas.
  • Ability to work with complex scientific/biomarker data or ramp quickly.
  • Experience handling PHI-adjacent clinical data under contractor policies.
  • Willingness to work with legacy R or PySpark code to extract rules.

Responsibilities

  • Build and maintain orchestrated ingestion pipelines for genomics, proteomics, and other data sources.
  • Develop layered transformation models with test coverage and data-quality guardrails.
  • Implement clinical data ingestion and reconciliation against standards (SDTM/ADaM).
  • Deliver platform infrastructure: APIs, CI/CD, containers, observability, and performance tuning.
  • Migrate legacy logic to new platform implementations, codify rules.
  • Translate research requirements into durable data models and contracts.
  • Identify repetitive processes and automate workflows with guardrails and AI-assisted tooling.
  • Participate in design reviews and enforce robust patterns.
  • Collaborate with lead engineer to accelerate delivery through paired sessions and PR reviews.
  • Ensure reproducibility with CI on PRs and automated tests.

Skills

AI-native engineering
Python
SQL
dbt
Dagster
Airflow
Data quality

Education

Bachelor's or Master's in CS/Data Eng/Bioinformatics

Tools

R
PySpark
Docker/ECS

Job description

Job Summary

We are seeking a contract Senior Translational Data and AI Engineer to help modernize how biomarker and clinical data are ingested, transformed, and delivered within Computational Discovery. The group is consolidating a fragmented set of ETL processes onto a modern, AWS-native lakehouse platform. Core infrastructure is already in place; this role adds the engineering capacity to bring real scientific and clinical data through production pipelines across multiple studies, and to serve it reliably to downstream analytics, AI, and visualization consumers.
The successful candidate will work as a hands-on technical partner alongside the lead engineer, contributing across the full stack from source ingestion through curated delivery. This is a role that pairs strong data-quality instincts with an AI-native way of working: not just using AI coding agents, but building scalable, guarded workflows around them. Top-level architecture and design decisions remain owned by the lead engineer; we are looking for a strong senior engineer who executes independently within that direction and is comfortable in scientific data domains - or ramps into them quickly.

Key Responsibilities
  • Build and maintain orchestrated ingestion pipelines for external genomics, proteomics, and other assay data sources, including source IO, table-format writers, and row-level reconciliation.
  • Develop and harden layered transformation models (staging, intermediate, and mart) with real-data test coverage, data-quality guardrails, and reusable, consolidated logic.
  • Implement clinical data ingestion and reconciliation paths against recognized standards (e.g., SDTM, ADaM), including subject and entity resolution.
  • Deliver supporting platform infrastructure: service APIs, CI/CD pipelines, containerized deployments, observability instrumentation, and data-warehouse performance tuning.
  • Extract transformation logic and business rules from legacy analytical code (e.g., R, PySpark) and reconcile them against new platform implementations.
  • Translate scientific and biomarker requirements from research and bioinformatics partners into durable data models and published data contracts.
  • Identify repetitive processes and convert them into automated workflows, guardrails, or reusable tooling — including AI-assisted workflows that make future work faster.
  • Participate in adversarial design and code reviews, identifying edge cases and pushing back on suboptimal patterns.
  • Collaborate with the lead engineer on design decisions and support delivery velocity through paired working sessions and PR reviews.
  • Ensure all work meets reproducibility standards: CI on every PR, automated tests, and no ad-hoc notebook-based production processes.
Minimum Qualifications
  • AI-native engineering practice: demonstrated experience building systems and workflows around AI coding agents (Claude Code, Cursor, Codex, or equivalent) — not just prompting them. You recognize when a repeated process should become an automated pipeline, when agent output needs guardrails, and when to build infrastructure that makes future work faster. Surface-level tool usage is insufficient.
  • Education: Bachelor's or master's degree in Computer Science, Data Engineering, Bioinformatics, or a related field.
  • Experience: 4+ years of professional experience in data engineering with shipped production pipelines on AWS (S3, ECS/Fargate, Redshift or equivalent MPP).
  • Strong proficiency in Python and SQL with working knowledge of modern data engineering libraries.
  • Solid, hands-on experience with dbt and a workflow orchestration tool (Dagster, Airflow, or Prefect).
  • Data quality instinct: track record of catching silent failures, questioning data correctness assumptions, and noticing lossy joins or incomplete deliveries.
  • Working understanding of lakehouse architecture patterns, ETL processes, and schema design for complex multi-modal datasets.
  • Comfort working with scientific, biomarker, or other complex domain data — or a demonstrated ability to ramp on unfamiliar scientific domains quickly.
  • Ability to handle PHI-adjacent clinical data under contractor policy (background check, compliance training, VPN access).
  • Willingness to work within legacy codebases (R, PySpark) to extract business rules and validate new implementations.
  • Excellent communication skills and ability to work in an embedded pair model with tight feedback loops.
Preferred Qualifications
  • Experience building or maintaining tooling around AI coding agents (custom commands, subagents, evals, or guardrails) rather than only consuming them.
  • Direct experience with Apache Iceberg, AWS Glue Catalog, or lakehouse table formats.
  • Preferred fluency reading genomic data (VAF, HGVS nomenclature, VCFs, CNV/fusion semantics).
  • Familiarity with clinical data standards including SDTM, ADaM, and CDISC.
  • Pharma, clinical research, or life sciences background.
  • Experience with containerization (Docker/ECS) and infrastructure-as-code (CloudFormation).
  • Proficiency in R for interoperability with bioinformatics teams.

Kaztronix is an equal opportunity employer and does not discriminate on the basis of race, color, national origin, sex, age, religion, disability, veteran status or any other consideration made unlawful by federal, state or local laws. In addition, all human resource actions in such areas as compensation, employee benefits, transfers, layoffs, training and development are to be administered objectively, without regard to race, color, religion, age, sex, national origin, disability, veteran status or any other consideration made unlawful by federal, state or local laws.

By applying to the position, you acknowledge that your information will be used by Kaztronix in processing your application.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Translational Data & AI Engineer - AI-Native Pipelines
Senior Translational Data & AI Engineer - AI-Native Pipelines

Creative Solutions Services, LLC • Wilmington (DE)

On-site
USD 140,000 - 180,000
Data Engineer
Data Engineer

BioAgilytix • North Carolina

On-site
USD 120,000 - 160,000
Medical Insurance (HDHP with HSA)
Dental Insurance
Vision Insurance
+4
Data Engineer
Data Engineer

BioAgilytix Labs, LLC • Durham (NC)

On-site
USD 110,000 - 165,000
Medical Insurance
401k Match
Paid Time Off
+1
Data Engineer
Data Engineer

BioAgilytix • Durham (NC)

On-site
USD 110,000 - 165,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Software Engineer Senior
Software Engineer Senior

Nationwide Children's Hospital • Columbus (OH)

On-site
USD 120,000 - 160,000
Data Engineer
Data Engineer

Axle • Rockville (MD)

On-site
USD 85,000 - 110,000
Paid Time Off
401K match up to 5%
Educational Benefits for Career Growth
+1
Senior Data Engineer
Senior Data Engineer

Tiger Analytics • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 190,000
Data Architect - Austin
Data Architect - Austin

Biorce • Austin (TX)

On-site
USD 140,000 - 210,000
Hybrid work model
MacBook and AI toolchain
Sr. Data Scientist, Clinical Data Solutions | Onsite San Diego HQ
Sr. Data Scientist, Clinical Data Solutions | Onsite San Diego HQ

Neurocrine Biosciences • San Diego (CA)

On-site
USD 103,000 - 141,000
Senior AI Data Scientist I
Senior AI Data Scientist I

Exelixis • Alameda (CA)

On-site
USD 143,000 - 203,000
401(k) with company contributions
Health, dental, vision
Life and disability insurance
+2