NLP Post-doc / Engineer for Information Mining in Historical Data (French 3rd Republic)

Inria

Paris

Sur place

EUR 42 000 - 55 000

Plein temps

Il y a 34 heures
Soyez parmi les premiers à postuler
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Avantages offerts par ce poste

Remboursement partiel des transports
Congés annuels importants (7 semaines)

Résumé du poste

Inria, au sein du projet ALMAnaCH, recherche un post-doctorant/ingénieur NLP pour l’exploration et l’extraction d’informations à partir de données historiques de la Troisième République française. Vous travaillerez en collaboration avec des historiens et des chercheurs du domaine, en vue de déployer des méthodes robustes d’analyse et de récupération sur un corpus historique volumineux.

Le candidat idéal maîtrise Python, des techniques de formation et évaluation de modèles NLP et a une affinité

Qualifications

  • Développement de recherches sur le sujet, positionnement par rapport aux travaux et validation par des expériences et analyses.
  • Conception d’une solution déployable pour l’information retrieval post-lexical.
  • Interaction avec les historiens et les chercheurs du lot NLP.
  • Présentation du travail en interne et en externe (conférences/journaux).
  • Échange avec les collègues sur les sujets NLP.

Responsabilités

  • Conduire des recherches sur le sujet, développer des idées nouvelles et valider les méthodes.
  • Produire une solution exploitable pour l’extraction d’information post-lexicale.
  • Collaborer avec les historiens et l’équipe ALMAnaCH sur les problématiques NLP.
  • Présenter les résultats en interne et lors de conférences/ateliers.
  • Échanger avec les collègues sur les problématiques NLP et bibliographie associée.

Connaissances

Python
NLP model training

Formation

PhD ou Master en domaine pertinent

Description du poste

NLP Post-doc / Engineer for Information Mining in Historical Data (French 3rd Republic)

This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on:

  • Development of an annotation interface to support the creation of thematic datasets across the corpus of parliamentary debates and, potentially, the press.
  • Development of an easily deployable approach for robust information retrieval across the corpus, targeting topics whose vocabulary may have evolved over time and diverged from contemporary French.
  • Evaluation of information retrieval methods for topic-based search on a historically situated corpus (including RAG).

The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA)

They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.

Participation to national meetings and national/international conference are to be expected.

The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves:

  • Setting up an annotation workflow (true/false positive labeling of keyword occurrences in context) in collaboration with historians;
  • Training and evaluating small-scale classifiers — ranging lightweight models to larger pretrained models — capable of distinguishing relevant from irrelevant occurrences of ambiguous or polysemous terms in diachrony (e.g., distinguishing an "anti-parliamentary" use of réforme de l'État from a routine administrative reform);
  • Integrating active learning so that model performance improves iteratively as historians annotate, minimizing labeling effort while maximizing corpus quality;
  • Prioritizing model deployability: given the scale of the corpus (at least dozens of millions of tokens across noisy OCR output) and the need for the tool to run efficiently and reproducibly within a web interface used directly by non-specialist historians, the research should focus on enabling this on lightway models: fast at inference, and easy to retrain/update rather than relying on large LLM inference at scale;
  • Benchmarking against and complementing RAG-based exploration (T5.5), providing a transparent, low-cost alternative for corpus-scoping that historians can audit and reproduce before moving to more exploratory or generative tasks.

The research component focuses on efficient, low-resource sequence classification for historical/noisy text — including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline.

Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., Villemonte de la Clergerie, É., Seddah, D., & Sagot, B. (2020). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.645

van Strien, D., Beelen, K., Coll Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the impact of OCR quality on downstream NLP tasks. In A. Rocha, L. Steels, & J. van den Herik (Eds.), Proceedings of the 12th International Conference on Agents and Artificial Intelligence: ICAART 2020 (Vol. 1, pp. 484–496). SCITEPRESS. https://doi.org/10.5220/0009169004840496

Bian, D., Puren, M., & Cafiero, F. (2026). How to efficiently explore noisy historical data? Leveraging corpus pre-targeting to enhance graph-based RAG. In D. Alves, Y. Bizzoni, S. Degaetano-Ortlieb, A. Kazantseva, J. Pagel, & S. Szpakowicz (Eds.), Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities, and Literature 2026 (pp. 241–250). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.latechclfl-1.23

Clerice, T. (2024). Detecting sexual content at the sentence level in first millennium Latin texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4772–4783). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.427/

The main activities of the applicant will include:

  • carrying out research on the topic outlined above, both in the development of new ideas, positioning with respect to related work and validation of the methodology via experiments and analysis
  • Producing a solution that will be deployable for post-lexical information retrieval
  • Interact with the project’s historians as well as the researchers from the work package on natural language processing
  • the presentation of work both internally to colleagues and externally in the form of conference/journal/workshop papers
  • interacting and exchanging with colleagues on NLP topics

Technical skills and level required :

  • Python
  • NLP Model Training and evaluation, beyond just LLM

Languages :

  • French (Reading understanding to interact with the data)

Relational skills :

  • Good organizational skills.
  • Good interpersonal skills.

Additional skills considered an asset:

  • Knowledge of the French 3rd Republic or similar political systems.
  • Interefest in information retrieval
Avantages
  • Partial reimbursement of public transport costs
  • Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
  • Possibility of teleworking and flexible organization of working hours
  • Professional equipment available (videoconferencing, loan of computer equipment, etc.)
  • Social, cultural and sports events and activities
  • Feeling comfortable in an interdisciplinary environment, as well as a willingness to learn and listen, are essential qualities for success in this role.
  • Interest in open science issues.
  • A PhD thesis or master's dissertation focusing on historical data or on information retrieval is an asset.
  • Interest in small models, beyond LLMs
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Post-doctoral researcher in mechanistic interpretability and diversity of representations
Post-doctoral researcher in mechanistic interpretability and diversity of representations

IMT Atlantique • France

Sur place
EUR 36 000 - 45 000
Partial working from home
Public transport paid for
PhD Position F/M Pretrained models of multimodal neuroimaging for predicting individual cognition
PhD Position F/M Pretrained models of multimodal neuroimaging for predicting individual cognition

Inria • Palaiseau

Hybride
EUR 18 000 - 30 000
Remboursement des frais de transport
7 semaines de congés + RTT
Télétravail possible et organisation d
+2
Ingénieur scientifique contractuel / Production, traitement et analyse des données / Langue, parole et audio / Bac + 5 ou équivalent
Ingénieur scientifique contractuel / Production, traitement et analyse des données / Langue, parole et audio / Bac + 5 ou équivalent

Inria • Paris

Hybride
EUR 52 000 - 72 000
Restauration subventionnée
Transports publics remboursés
Congés annuels 7 semaines + RTT
+3
Post-Doctoral Research Visit F/M Trait-Based Species Identification via Knowledge Extraction an[...]
Post-Doctoral Research Visit F/M Trait-Based Species Identification via Knowledge Extraction an[...]

Inria • Montpellier

Sur place
EUR 42 000 - 54 000
Transport refund
Annual leave
Telework after 6 months
+2
PhD Position F/M Spatio-temporal analysis of remote sensing data at large scales
PhD Position F/M Spatio-temporal analysis of remote sensing data at large scales

Inria • Valbonne

Sur place
EUR 23 000 - 28 000
Remboursement partiel des frais de bus
Congés annuels 7 semaines + RTT
Télétravail possible
+4
Phd Position F - M Visualization Of The Plausibility And Bias For Data Resources Used In a Geographic Digital Twin H/F
Phd Position F - M Visualization Of The Plausibility And Bias For Data Resources Used In a Geographic Digital Twin H/F

INRIA • Villé

Sur place
EUR 22 000 - 30 000
Subsidized meals
Public transport reimbursement
Leave + RTT
+5
Chercheur Géographe philo-rudologue / Researcher geographer philo-rudologist
Chercheur Géographe philo-rudologue / Researcher geographer philo-rudologist

Faculté de Droit, Sciences Économiques & Gestion | Le Mans Université • Le Mans

Sur place
EUR 42 000 - 52 000
Transport 75% couverts
Mutuelle 50% prise en charge
Prevoyance optionnelle
+1
Research Master Internship: New Algorithms for Automated Room Acoustic Diagnosis
Research Master Internship: New Algorithms for Automated Room Acoustic Diagnosis

Inria • Strasbourg

Hybride
EUR 6 700 - 10 000
Partial transport reimbursement
6 months teleworking possibility
Flexible working hours
+1
Research engineer in AI security/fairness for specifying and implementing regulatory compliance tests for the EU AI Act
Research engineer in AI security/fairness for specifying and implementing regulatory compliance tests for the EU AI Act

Inria • Villeurbanne

Hybride
EUR 27 000 - 33 000
Partial transport reimbursement
Teleworking (90 days/yr)
Professional training
+3
Conception et mise en oeuvre d'un protocole d'entraînement d'un LLM médical français
Conception et mise en oeuvre d'un protocole d'entraînement d'un LLM médical français

Inria • Paris

Hybride
EUR 33 000 - 42 000
Restauration subventionnée
Transports publics remboursés
Congés: 7 semaines + 10 RTT
+3