Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à être envoyés.
Inria Rennes propose une thèse doctorante sur RobDNA: récupération robuste de données dans le stockage ADN. Le sujet combine compression et correction d'erreurs pour des lectures nanopore, en lien avec la transformation en alphabet quaternaire et le décodage non-binaire.
Le candidat développera des algorithmes et évaluera des performances sur des ensembles de données réels et simulés, en participant à la diffusion scientifique et à l'intégration dans l'écosystème Inria.
Fonction : Doctorant
The Inria Centre at Rennes University is one of Inria's nine centres and has more than thirty research teams. The Inria Centre is a major and recognized player in the field of digital sciences. It is at the heart of a rich R&D and innovation ecosystem: highly innovative PMEs, large industrial groups, competitiveness clusters, research and higher education players, laboratories of excellence, technological research institute, etc.
Context The volume of data generated worldwide is projected to approach 180 zettabytes (ZB) per year by 2025 [1]. However, current storage technologies face significant limitations in scaling sustainably to such volumes. One promising solution to address these challenges is DNA-based data storage, which offers several advantages, including extremely high data density, long-term retention, and low energy consumption [2].
From a density perspective, DNA can theoretically store up to 10 terabytes per mm³, which would allow all data generated throughout human history to be stored within a cube of approximately 30 cm per side [3]. In terms of retention, DNA can remain readable for centuries under suitable conditions, whereas conventional storage media typically degrade within decades [3]. Furthermore, DNA storage is energy-efficient, as it can be preserved at ambient temperature provided it is protected from light and humidity.
Goal The goal of the project is to develop an algorithm to allow robust retrieval of data in the context of DNA-based data storage.
Challenges and envisaged approach Despite its potential, making DNA a practical and efficient storage medium requires overcoming several key challenges:
(i) Data transformation: converting digital data into a quaternary alphabet (A, C, G, T).
(ii) DNA synthesis: writing data through the physical synthesis of DNA strands.
(iii) DNA sequencing: reading the stored data by sequencing DNA.
(iv) Data retrieval: reconstructing the original digital data from the sequenced symbols.
This PhD project focuses on the first and fourth challenges, by developing joint compression and error-correction algorithms that are robust to sequencing errors arising during step (iii).
Efficient DNA storage critically depends on fast sequencing technologies, which often come at the cost of increased error rates. For example, nanopore sequencing, developed by Oxford Nanopore Technologies (ONT), enables real-time analysis but introduces relatively high error rates [4,5]. Unlike traditional sequencing technologies, nanopore sequencing produces not only substitution errors but also insertion and deletion errors. Deletions are particularly challenging, as they differ from erasure errors where the position of missing data is known
e.g., packet losses in digital communications). In the case of deletions, neither the existence nor the location of the missing symbols is known, significantly complicating error correction.
Two main approaches have been proposed in the literature to address these errors. The first consists of avoiding error-prone patterns by designing constrained DNA sequences, such as limiting homopolymers or enforcing balanced GC content [6]. The second approach embraces sequencing errors and focuses on correcting them using coding techniques [7]. These strategies are typically considered contradictory, as one seeks to prevent errors while the other assumes their presence [8].
In this project, we propose to explore an alternative route that combines both strategies: avoiding the majority of sequencing errors while correcting the remaining ones. This will be achieved by jointly structuring the compressed DNA stream and designing error-correction mechanisms tailored to nanopore sequencing. In particular, we will exploit the properties of the quaternary alphabet and the interaction between compression and error-correction algorithms. The work will build upon concepts from non-binary channel coding [10,11] and DNA-specific transcoding techniques [6], while also integrating efficient post-sequencing data retrieval methods such as those proposed in [9].
[1] David Reinsel-John Gantz-John Rydning, John Reinsel, and John Gantz. “The digitization of the world from edge to core.Framingham: International Data Corporation”, 16:1–28, 2018.
[2] Luis Ceze, Jeff Nivala, and Karin Strauss. “Molecular digital data storage using DNA”. NatureReviews Genetics, 20(8):456–466, 2019.
[3] Victor Zhirnov, Reza M Zadegan, Gurtej S Sandhu, George M Church, and William L Hughes. “Nucleic acid memory”. Nature materials, 15(4):366–370, 2016.
[4] Delahaye, Clara, and Jacques Nicolas. “Nanopore MinION Long Read Sequencer: An Overview of Its Error Landscape,” November 23, 2020. https://hal.inria.fr/hal-03123133.
[5] ———. “Sequencing DNA with Nanopores: Troubles and Biases.” PLoS ONE, October 1, 202.
[6] S. Al Sayyed and A. Roumy and T. Maugey. ``Efficient constraining of transcoding for DNA-based data storage'', IEEE International Conference on Image Processing (ICIP), 2025.
[7] R. Khabbaz, M. Antonini, S. Kas Hanna, Marker Guess & Check Plus (MGC+): An Efficient Short Blocklength Code for Random Edit Errors, International Symposium on Topics in Coding (ISTC), 2025.
[8] F. Weindel, A. L. Gimpel, R. N. Grass and R. Heckel, ”Embracing errors is more effective than avoiding them through constrained coding for DNA data storage,” Allerton Conference on Communication, Control, and Computing 2023.
[10] M. C. Davey and D. J. C. Mackay, “Low density parity check codes over GF(q)”, Information Theory Workshop, 1998, pp. 70-71.
[11] D. Declercq and M. Fossorier, “Decoding algorithms for nonbinary LDPC codes over GF(q)”, IEEE Trans. on Commun., vol. 55, pp. 633- 643, April 2007
Candidate profile The candidate should have