Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.
Inria, au sein du projet ALMAnaCH, recherche un post-doctorant/ingénieur NLP pour l’exploration et l’extraction d’informations à partir de données historiques de la Troisième République française. Vous travaillerez en collaboration avec des historiens et des chercheurs du domaine, en vue de déployer des méthodes robustes d’analyse et de récupération sur un corpus historique volumineux.
Le candidat idéal maîtrise Python, des techniques de formation et évaluation de modèles NLP et a une affinité
This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on:
The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA)
They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.
Participation to national meetings and national/international conference are to be expected.
The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves:
The research component focuses on efficient, low-resource sequence classification for historical/noisy text — including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline.
Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., Villemonte de la Clergerie, É., Seddah, D., & Sagot, B. (2020). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.645
van Strien, D., Beelen, K., Coll Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the impact of OCR quality on downstream NLP tasks. In A. Rocha, L. Steels, & J. van den Herik (Eds.), Proceedings of the 12th International Conference on Agents and Artificial Intelligence: ICAART 2020 (Vol. 1, pp. 484–496). SCITEPRESS. https://doi.org/10.5220/0009169004840496
Bian, D., Puren, M., & Cafiero, F. (2026). How to efficiently explore noisy historical data? Leveraging corpus pre-targeting to enhance graph-based RAG. In D. Alves, Y. Bizzoni, S. Degaetano-Ortlieb, A. Kazantseva, J. Pagel, & S. Szpakowicz (Eds.), Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities, and Literature 2026 (pp. 241–250). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.latechclfl-1.23
Clerice, T. (2024). Detecting sexual content at the sentence level in first millennium Latin texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4772–4783). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.427/
The main activities of the applicant will include:
Technical skills and level required :
Languages :
Relational skills :
Additional skills considered an asset: