N’envoyez pas de CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.
Rivercell in Paris is seeking a Bioinformatics Data Engineer to build and operate data pipelines used to train and evaluate AI-driven cell models. You will process sequencing and imaging data, ensure versioned, analysis-ready datasets, and enable reproducible experiments.
You will collaborate with ML scientists, biologists, chemists and engineers on an ambitious biotech startup project, with an emphasis on scalable pipelines, cloud/HPC infrastructure, and rigorous data management.
We are looking for a Bioinformatics Data Engineer to join our team and take charge of all the data used to train and evaluate our AI virtual cell models: processing, harmonization, and versioning.
In your daily work, you will write and run the bioinformatics pipelines that turn raw sequencing and imaging data into versioned, analysis-ready datasets.
You will work hand-in-hand with our ML scientists, biologists, chemists and engineers. You will be an integral part of an ambitious and exciting French startup project from its inception.
Build our omics pipelines, starting with single-cell, from FASTQ to analysis-ready matrices: alignment and quantification, CRISPR guide assignment, cell calling, ambient RNA and doublet handling, and per-run quality control
Build the imaging pipelines for live-cell brightfield and fluorescence time-lapse data: illumination correction, segmentation, tracking, and feature and embedding extraction
Internalize and harmonize relevant publicly available datasets with Rivercell's own data
Set up dataset versioning, lineage, and release management, so that every dataset behind a model, a paper, or a benchmark can be traced to its raw data and pipeline version and every experiment is reproducible
Build the tooling that generates and audits contamination-controlled train/test splits for our benchmarks
Run the data infrastructure: cloud storage and compute for tens of terabytes, cost control, and data loading fast enough to keep GPUs busy
Build the data side of our closed lab-in-a-loop, in which models propose experiments, the lab runs them, and the results return as retraining-ready data
Opportunity to work on a cutting-edge project in the field of biotech
Opportunity to join an ambitious team and impactful startup among the first employees
Competitive salary and equity package
Exciting opportunities for personal and professional growth within the team
You have an MSc or PhD in bioinformatics, computational biology, computer science, or a related field You have 3+ years of experience building and running bioinformatics pipelines in production, in industry, a core facility, or a large consortium You have processed single-cell RNA-seq data from raw reads (Cell Ranger, STARsolo, kallisto/bustools, alevin-fry, or equivalent) and work comfortably in the AnnData / scverse ecosystem You write strong Python and apply sound engineering practice: workflow managers, containers, tests, CI, and documentation You have experience with data versioning on large datasets, with harmonizing heterogeneous datasets, and with cloud (AWS or GCP) or HPC infrastructure You have experience with image processing at scale, ideally microscopy: segmentation, tracking, and feature or embedding extraction (Cellpose, StarDist, CellProfiler, or deep-learning models) You have experience with high-performance data management for model training: array and columnar formats (Zarr, TileDB-SOMA, Parquet, OME-Zarr), sharding and streaming from object storage, and PyTorch data loaders that keep GPUs busy An autonomous, committed, enthusiastic, fun to work with person with a "can-do", creative, and challenging mindset You have excellent communication skills and are fluent in English. Professional proficiency in French is helpful but not required.