PhD Position F/M RobDNA:Robust data retrieval for DNA-based data storage

Inria

Rennes

Sur place

EUR 21 000 - 26 000

Plein temps

Il y a 12 jours
Générateur de candidature

Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à être envoyés.

Passez les filtres ATS

Avantages offerts par ce poste

Remboursement transport
Congés annuels + RTT
Télétravail après 6 mois
Equipement professionnel
Événements sociaux et sportifs

Résumé du poste

Inria Rennes propose une thèse doctorante sur RobDNA: récupération robuste de données dans le stockage ADN. Le sujet combine compression et correction d'erreurs pour des lectures nanopore, en lien avec la transformation en alphabet quaternaire et le décodage non-binaire.

Le candidat développera des algorithmes et évaluera des performances sur des ensembles de données réels et simulés, en participant à la diffusion scientifique et à l'intégration dans l'écosystème Inria.

Qualifications

  • Solide background en traitement d'image et traitement du signal.
  • Connaissances en optimisation et programmation.
  • Notions de codage source et théorie de l'information appréciées.

Responsabilités

  • Développer des algorithmes de compression et de correction d'erreurs robustes au bruit de lecture nanopore.
  • Intégrer les techniques de transcoding et évaluer les performances sur des jeux de données simulés et réels.
  • Participer à la rédaction de publications et à la supervision des expérimentations.

Connaissances

Traitement d'image
Traitement du signal
Optimisation
Programmation

Description du poste

PhD Position F/M RobDNA:Robust data retrieval for DNA-based data storage

Fonction : Doctorant

The Inria Centre at Rennes University is one of Inria's nine centres and has more than thirty research teams. The Inria Centre is a major and recognized player in the field of digital sciences. It is at the heart of a rich R&D and innovation ecosystem: highly innovative PMEs, large industrial groups, competitiveness clusters, research and higher education players, laboratories of excellence, technological research institute, etc.

Context The volume of data generated worldwide is projected to approach 180 zettabytes (ZB) per year by 2025 [1]. However, current storage technologies face significant limitations in scaling sustainably to such volumes. One promising solution to address these challenges is DNA-based data storage, which offers several advantages, including extremely high data density, long-term retention, and low energy consumption [2].

From a density perspective, DNA can theoretically store up to 10 terabytes per mm³, which would allow all data generated throughout human history to be stored within a cube of approximately 30 cm per side [3]. In terms of retention, DNA can remain readable for centuries under suitable conditions, whereas conventional storage media typically degrade within decades [3]. Furthermore, DNA storage is energy-efficient, as it can be preserved at ambient temperature provided it is protected from light and humidity.

Goal The goal of the project is to develop an algorithm to allow robust retrieval of data in the context of DNA-based data storage.

Challenges and envisaged approach Despite its potential, making DNA a practical and efficient storage medium requires overcoming several key challenges:

(i) Data transformation: converting digital data into a quaternary alphabet (A, C, G, T).

(ii) DNA synthesis: writing data through the physical synthesis of DNA strands.

(iii) DNA sequencing: reading the stored data by sequencing DNA.

(iv) Data retrieval: reconstructing the original digital data from the sequenced symbols.

This PhD project focuses on the first and fourth challenges, by developing joint compression and error-correction algorithms that are robust to sequencing errors arising during step (iii).

Efficient DNA storage critically depends on fast sequencing technologies, which often come at the cost of increased error rates. For example, nanopore sequencing, developed by Oxford Nanopore Technologies (ONT), enables real-time analysis but introduces relatively high error rates [4,5]. Unlike traditional sequencing technologies, nanopore sequencing produces not only substitution errors but also insertion and deletion errors. Deletions are particularly challenging, as they differ from erasure errors where the position of missing data is known

e.g., packet losses in digital communications). In the case of deletions, neither the existence nor the location of the missing symbols is known, significantly complicating error correction.

Two main approaches have been proposed in the literature to address these errors. The first consists of avoiding error-prone patterns by designing constrained DNA sequences, such as limiting homopolymers or enforcing balanced GC content [6]. The second approach embraces sequencing errors and focuses on correcting them using coding techniques [7]. These strategies are typically considered contradictory, as one seeks to prevent errors while the other assumes their presence [8].

In this project, we propose to explore an alternative route that combines both strategies: avoiding the majority of sequencing errors while correcting the remaining ones. This will be achieved by jointly structuring the compressed DNA stream and designing error-correction mechanisms tailored to nanopore sequencing. In particular, we will exploit the properties of the quaternary alphabet and the interaction between compression and error-correction algorithms. The work will build upon concepts from non-binary channel coding [10,11] and DNA-specific transcoding techniques [6], while also integrating efficient post-sequencing data retrieval methods such as those proposed in [9].

[1] David Reinsel-John Gantz-John Rydning, John Reinsel, and John Gantz. “The digitization of the world from edge to core.Framingham: International Data Corporation”, 16:1–28, 2018.

[2] Luis Ceze, Jeff Nivala, and Karin Strauss. “Molecular digital data storage using DNA”. NatureReviews Genetics, 20(8):456–466, 2019.

[3] Victor Zhirnov, Reza M Zadegan, Gurtej S Sandhu, George M Church, and William L Hughes. “Nucleic acid memory”. Nature materials, 15(4):366–370, 2016.

[4] Delahaye, Clara, and Jacques Nicolas. “Nanopore MinION Long Read Sequencer: An Overview of Its Error Landscape,” November 23, 2020. https://hal.inria.fr/hal-03123133.

[5] ———. “Sequencing DNA with Nanopores: Troubles and Biases.” PLoS ONE, October 1, 202.

[6] S. Al Sayyed and A. Roumy and T. Maugey. ``Efficient constraining of transcoding for DNA-based data storage'', IEEE International Conference on Image Processing (ICIP), 2025.

[7] R. Khabbaz, M. Antonini, S. Kas Hanna, Marker Guess & Check Plus (MGC+): An Efficient Short Blocklength Code for Random Edit Errors, International Symposium on Topics in Coding (ISTC), 2025.

[8] F. Weindel, A. L. Gimpel, R. N. Grass and R. Heckel, ”Embracing errors is more effective than avoiding them through constrained coding for DNA data storage,” Allerton Conference on Communication, Control, and Computing 2023.

[10] M. C. Davey and D. J. C. Mackay, “Low density parity check codes over GF(q)”, Information Theory Workshop, 1998, pp. 70-71.

[11] D. Declercq and M. Fossorier, “Decoding algorithms for nonbinary LDPC codes over GF(q)”, IEEE Trans. on Commun., vol. 55, pp. 633- 643, April 2007

Avantages
  • Partial reimbursement of public transport costs
  • Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
  • Possibility of teleworking (after 6 months of employment) and flexible organization of working hours
  • Professional equipment available (videoconferencing, loan of computer equipment, etc.)
  • Social, cultural and sports events and activities

Candidate profile The candidate should have

  • strong background in image/signal processing, optimization and programming,
  • notions of source coding, information theory would be appreciated.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

PhD in DNA Data Storage: Robust Data Retrieval & Coding
PhD in DNA Data Storage: Robust Data Retrieval & Coding

Inria • Rennes

Sur place
EUR 21 000 - 26 000
Remboursement transport
Congés annuels + RTT
Télétravail après 6 mois
+2
NGS Specialist
NGS Specialist

Biomemory • Paris

Sur place
EUR 70 000 - 90 000
Flexible working hours
Additional benefits beyond French regulations
Software Engineer
Software Engineer

Biomemory • Paris

Sur place
EUR 60 000 - 90 000
Flexible work hours
Employee well-being benefits
R&D-Biotech engineer
R&D-Biotech engineer

BioMol • France

Sur place
EUR 60 000 - 90 000
M2 internship: Optical sequencing of digital polymers for molecular data storage
M2 internship: Optical sequencing of digital polymers for molecular data storage

France-BioImaging • Montpellier

Sur place
M2 internship: Optical sequencing of digital polymers for molecular data storage
M2 internship: Optical sequencing of digital polymers for molecular data storage

France - BioImaging • Marseille

Sur place
EUR 7 800 - 12 000
Close supervision by expert
Access to state-of-the-art equipment
Personal and professional development training
R&D Automation Specialist
R&D Automation Specialist

Biomemory • Paris

Sur place
EUR 35 000 - 50 000
PhD Position F/M Spatio-temporal analysis of remote sensing data at large scales
PhD Position F/M Spatio-temporal analysis of remote sensing data at large scales

Inria • Valbonne

Sur place
EUR 23 000 - 28 000
Remboursement partiel des frais de bus
Congés annuels 7 semaines + RTT
Télétravail possible
+4
Postdoctoral Researcher in Bioconjugation and Oligonucleotide Chemistry (M/F)
Postdoctoral Researcher in Bioconjugation and Oligonucleotide Chemistry (M/F)

CNRS - National Center for Scientific Research • France

Hybride
EUR 42 000 - 52 000
PhD Position F/M Frugal Distributed Training with Volatile Resources
PhD Position F/M Frugal Distributed Training with Volatile Resources

Inria • Valbonne

Sur place
Partial reimbursement of public transport costs
7 weeks of annual leave + 10 extra days off
Possibility of teleworking
+3