Senior Data Engineer — Medical AI Data Platform (m/f/d)

Cancilico

Dresden

Hybrid

EUR 70.000 - 95.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Salary range €70,000-€95,000
30 days paid leave
40-hour weeks
Hybrid in Dresden
Remote within EU time zones
Choice of OS
English as working language

Zusammenfassung

Cancilico, a Dresden-based health-tech startup, seeks a Senior Data Engineer to own and evolve the Medical AI Data Platform. You’ll design robust data registries, ensure provenance, and enable reproducible training datasets for our AI models.

You will write production Python and SQL, build data pipelines from hospital sources, and collaborate with AI and clinical teams. Hybrid in Dresden with EU remote options reflects our flexible approach.

Qualifikationen

  • Strong hands-on software and data-engineering experience writing production-quality Python and SQL and operating data systems.
  • Strong Python and SQL, including practical PostgreSQL data modelling, query design, and schema evolution.
  • Experience building idempotent and recoverable pipelines with validation, retries, observability, and failure handling.
  • Understanding data contracts, provenance, versioning, and trade-offs between relational metadata and large assets in object storage.
  • Strong testing and debugging discipline to guard against data loss, duplication, leakage, and partial state.
  • Product-minded ownership turning AI/clinical needs into a coherent platform.
  • Familiar with modern AI coding tools in development workflows.

Aufgaben

  • Own the application and data architecture of the internal data registry from raw assets to curated datasets for training and evaluation.
  • Build reliable ingestion pipelines from hospitals, scanners, and annotation systems with safe contracts and validation at boundaries.
  • Evolve PostgreSQL data models and interfaces for images, annotations, patients, samples, and provenance; design history-preserving migrations.
  • Make data quality operational via identity resolution and deduplication workflows with checksum and metadata controls; detect drift and leakage.
  • Publish reproducible datasets linking database state, object-store artifacts, code, and manifests for traceability of training runs.
  • Prepare multimodal data support including clinical values and additional imaging modalities without overcomplicating the core model.
  • Operate the platform as a production system with observability, recovery, access controls, and automated tests; collaborate with infra teams.

Kenntnisse

Python
SQL
PostgreSQL
Data pipelines
Data governance
Testing & debugging
MLOps tooling

Tools

DVC
Prefect
S3-compatible storage
SQLAlchemy
Alembic
Kubernetes

Jobbeschreibung

Senior Data Engineer - Medical AI Data Platform (m/f/d)

Full time · 40h · Dresden (hybrid) · remote within EU possible · €70,000-95,000

Why this role exists

Cancilico is a Dresden-based, seed-funded health-tech company, spun out of TU Dresden and University Hospital Dresden. Our first product, MyeloAID, uses AI to automate the morphological differentiation of bone marrow smears. No one else in the EU is doing AI-driven bone marrow analysis at this depth, and we are expanding into new diagnostic verticals and, over time, multimodal models that combine morphology with clinical, laboratory, and molecular data.

Those ambitions depend on more than training better models. We need to know where every image, annotation, and clinical attribute came from; how it was transformed; whether two records represent an exact sample or patient; and which exact data supported an experiment, evaluation, or release.

We have built the foundations of that system: an internal data registry combining PostgreSQL, object storage, DVC, and orchestrated processing pipelines. It manages raw whole-slide and microscope images, metadata history, patient and sample identity, deduplication, protected holdouts, and curated datasets for model development.

Your job is to own and evolve this platform into the dependable data backbone for our computer‑vision products and future multimodal AI. You will make complex medical data discoverable, reproducible, secure, and usable by the people building and validating our models.

What you’ll do

This is a deeply hands‑on engineering role. You will personally write, test, review, deploy, debug, and maintain production Python and SQL—not only design the system or coordinate its implementation.

  • Own the application and data architecture of our internal data registry, from immutable raw assets through enriched and deduplicated records to curated, versioned datasets for training and evaluation.
  • Build reliable ingestion pipelines for data from hospitals, scanners, microscope cameras, annotation systems, and product workflows. Define contracts and validation at each boundary so partial uploads, retries, and malformed metadata fail safely and visibly.
  • Evolve our PostgreSQL data model and interfaces for images, annotations, patients, samples, clinical metadata, dataset membership, and provenance. Design migrations that preserve history and keep a live system trustworthy as its schema changes.
  • Make data quality operational. Own the platform workflows around identity resolution and deduplication, including checksum, metadata, and perceptual-hash controls. Work with the AI/Data Science team, who own the advanced image‑similarity models, to integrate their methods into those workflows. Detect inconsistencies and drift, protect patient‑level holdouts from leakage, and give operators clear tools to inspect and resolve ambiguous cases.
  • Publish reproducible datasets that AI engineers can consume confidently. Connect database state, object‑store artifacts, code, and manifests so a training run or released model can always be traced back to the exact data it used.
  • Prepare the platform for multimodal data: clinical and laboratory values, molecular findings, additional imaging modalities, and changing annotation taxonomies, without turning the core model into an unmaintainable collection of one‑off fields.
  • Operate the platform as a production system. Improve observability, reconciliation, recovery, access boundaries, documentation, and automated tests while working with our infrastructure engineers on the services underneath it.
  • Work with QM/RA to translate data‑governance and change‑control needs into practical technical controls and evidence. You will contribute the traceability and validation records behind compliance; you will not be expected to own regulatory submissions yourself.

You will collaborate closely with AI engineers, clinical experts, QM/RA, and infrastructure engineers. You will not be expected to own production model development, scientific performance evaluation, or the underlying Kubernetes and networking platform.

What we’re looking for
Must have
  • Strong hands‑on software and data‑engineering experience. You currently write production‑quality Python and SQL and have personally implemented, tested, deployed, debugged, and operated data systems—not only designed or managed them.
  • Strong Python and SQL, including practical PostgreSQL data modelling, query design, and schema evolution.
  • Experience building idempotent and recoverable batch or incremental pipelines, with deliberate validation, retries, observability, and failure handling.
  • A good understanding of data contracts, provenance, versioning, and the trade‑offs between relational metadata and large assets in object storage.
  • Strong testing and debugging discipline. You look for silent data loss, duplication, leakage, and partial state—not just whether the happy path completed.
  • Product‑minded ownership. You can turn the needs of AI engineers, clinical colleagues, and data operators into a platform that is coherent rather than a sequence of special cases.
  • Confidence using modern AI coding tools critically and effectively as part of your development workflow.
Nice to have
  • Experience building data systems in a regulated, safety‑critical, or highly controlled environment. Healthcare, life sciences, GxP, finance, aerospace, and similar backgrounds are all relevant, and we value this experience highly.
  • Experience with medical‑device development or verification, ISO 13485, IVDR/MDR, IEC 62304, computerized‑system validation, or clinical‑performance studies.
  • Experience with medical imaging, computer vision, or other large unstructured datasets. Whole-slide images, microscopy, DICOM, tiling, and image annotations are particularly relevant.
  • Hands‑on experience with S3‑compatible object storage, DVC or another data‑versioning system, and Prefect or a comparable workflow orchestrator.
  • Experience with entity resolution, record linkage, similarity search, or human‑in‑the‑loop deduplication and conflict review.
  • Familiarity with annotation workflows, label taxonomies, COCO‑style datasets, and inter‑rater disagreement.
  • Experience handling pseudonymised health data, role‑based access, retention rules, and auditable change histories.
  • Familiarity with how ML teams consume datasets for training, evaluation, and experiment tracking. You do not need to be a model researcher.

You do not need a medical background, but you should be genuinely curious about the clinical problem and comfortable learning the language needed to model it accurately.

Our stack & setup
  • Data platform: Python, PostgreSQL, SQLAlchemy, Alembic, S3‑compatible object storage, and DVC.
  • Pipelines and operations: Prefect for orchestration, with automated health, reconciliation, and data‑quality checks.
  • Downstream consumers: PyTorch training repositories, MLflow experiments, annotation tooling, and FastAPI‑based product services consume versioned data from the registry.
  • Engineering: GitHub for code, CI/CD, and tasks. The underlying platform is operated with our infrastructure engineers rather than being the primary responsibility of this role.

You’ll report to our CAIO and work alongside AI engineers, infrastructure and product engineers, QM/RA, and ~10 clinical annotation consultants (hematologists and medical technical assistants).

What we offer
  • Salary: €70,000-€95,000 depending on experience and fit.
  • 30 days paid leave, flexible core hours (10:00-15:00).
  • 40-hour weeks.
  • Pick your machine: Mac, Windows, or Ubuntu.
  • Flexible German benefits, negotiable based on your needs.
  • Working language is English. No German required.
  • Hematologists and medical technical assistants are colleagues you can walk over and ask.
  • Hybrid in Dresden preferred. Fully remote within EU time zones possible for the right candidate. You’ll need an existing right to work in the EU; we can’t sponsor relocation or visas at this stage.
  • We welcome applications regardless of background, parental status, disability, or neurodivergence.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

(Senior) Sales Manager – Hospital & Lab Diagnostics (m/f/d)
(Senior) Sales Manager – Hospital & Lab Diagnostics (m/f/d)

Cancilico • Dresden

Hybrid
EUR 75.000 - 95.000
30 days leave
Hybrid work
Travel opportunities
Regulatory Affairs Manager / PRRC (m/f/d)
Regulatory Affairs Manager / PRRC (m/f/d)

Cancilico • Dresden

Hybrid
EUR 55.000 - 70.000
Hybrid in Dresden
Flexible core hours (10:00–15:00)
Training in auditor qualification and,
Data Scientist / Data Engineer
Data Scientist / Data Engineer

United States Digital Space LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Competitive salary
Professional growth
Innovative work environment
+1
Senior Medical AI & Clinical Data Engineer (m/f/d)
Senior Medical AI & Clinical Data Engineer (m/f/d)

Goodly Technologies GmbH • München

Hybrid
EUR 110.000 - 140.000
Lead technical role
On-premises compute resources
Flexible hours, hybrid work
+2
Data Scientist / Data Engineer
Data Scientist / Data Engineer

Free resume • Berlin

Hybrid
EUR 53.000 - 59.000
Competitive salary
Comprehensive benefits
Remote work options
+2
Data Scientist / Data Engineer
Data Scientist / Data Engineer

Join • Berlin

Vor Ort
EUR 90.000 - 130.000
Competitive salary
Professional development
Innovative work environment
Senior Forward Deployed Platform Engineer (m/w/d)
Senior Forward Deployed Platform Engineer (m/w/d)

Recare Deutschland GmbH • Berlin

Remote
EUR 95.000 - 130.000
Edenred card
Extra vacation day
Remote-friendly & flexible hours
Senior Forward Deployed Platform Engineer (m/w/d)
Senior Forward Deployed Platform Engineer (m/w/d)

Recare • Deutschland

Hybrid
EUR 90.000 - 120.000
Edenred card
Extra vacation day
Remote-friendly with flexible hours
Senior AI Data Platform Engineer (m/w/d)
Senior AI Data Platform Engineer (m/w/d)

Recare Deutschland GmbH • Deutschland

Remote
EUR 90.000 - 140.000
Edenred card
Extra vacation day
Remote-friendly with flexible hours
Full Stack Engineer
Full Stack Engineer

Vincere • München

Vor Ort
EUR 90.000 - 120.000
Hybrid work in Munich