Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Inria ALMAnaCH, Paris, recrute un post-doctorant/NLP engineer pour concevoir et entraîner des modèles aidant les historiens à construire des sous-corpus à partir de vastes textes historiques numérisés.
Vous mettrez en place un flux d’annotation, évaluerez des classificateurs variés et favoriserez des solutions légères et déployables via une interface web destinée à des historiens non spécialistes.
This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on:
The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA). They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.
Participation to national meetings and national/international conference are to be expected.
The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves:
The research component focuses on efficient, low-resource sequence classification for historical/noisy text — including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline.
El Assadi, A., Muennighoff, N., & Lee, J. (2026). The embedder's dilemma: LLMs are better, but at what cost? arXiv. https://doi.org/10.48550/arXiv.2608.12875
Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., Villemonte de la Clergerie, É., Seddah, D., & Sagot, B. (2020). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.645
van Strien, D., Beelen, K., Coll Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the impact of OCR quality on downstream NLP tasks. In A. Rocha, L. Steels, & J. van den Herik (Eds.), Proceedings of the 12th International Conference on Agents and Artificial Intelligence: ICAART 2020 (Vol. 1, pp. 484–496). SCITEPRESS. https://doi.org/10.5220/0009169004840496
Luo, X., Shinnick, Z., Griesshaber, N., Wang, Y., Yu, J., Shi, F., Torr, P., & Lu, Y. (2026). Pretraining language models on historical text. arXiv. https://doi.org/10.48550/arXiv.2606.02991
Bian, D., Puren, M., & Cafiero, F. (2026). How to efficiently explore noisy historical data? Leveraging corpus pre-targeting to enhance graph-based RAG. In D. Alves, Y. Bizzoni, S. Degaetano-Ortlieb, A. Kazantseva, J. Pagel, & S. Szpakowicz (Eds.), Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature 2026 (pp. 241–250). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.latechclfl-1.23
Clerice, T. (2024). Detecting sexual content at the sentence level in first millennium Latin texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4772–4783). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.427/
Technical skills and level required :
Languages :
Relational skills :
Additional skills considered an asset: