Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction

Association Bernard Gregory

Villejuif

Sur place

EUR 20 000 - 23 000

Plein temps

14 jours+
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

Association Bernard Gregory propose un doctorat sur le sujet: Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction. Le contrat doctoral porte sur l’apprentissage multimodal et les modèles dynamiques à partir de données incomplètes, dans un cadre de Villejuif.

Le doctorant travaillera sur des représentations fiables et adaptatives selon les modalités disponibles. Le projet vise à développer des représentations latentes robustes, gérer l’irrégularité des

Qualifications

  • Diplôme exigé: Master en recherche ou diplôme d’ingénieur en informatique.
  • Forte expérience en apprentissage automatique et apprentissage profond.
  • Compétences fortes en programmation Python.
  • Capacité à comprendre et critiquer des papiers de recherche en ML.
  • Expérience en recherche via mémoire/thèse ou stage.
  • Expérience souhaitable en Self-Supervised Learning, Time-Series ou Probabilistic ML.

Responsabilités

  • Analyser et formuler les problématiques de multimodalité et de temps.
  • Développer des représentations latentes dynamiques et robustes.
  • Évaluer l’incertitude et assurer l’explicabilité des prédictions.
  • Publier et présenter les résultats dans des conférences et journaux.

Connaissances

Machine Learning
Deep Learning

Formation

Master’s degree (research-oriented) or an Engineering degree in computer science

Outils

Python

Description du poste

Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction

19/08/2026 Contrat doctoral

Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction

Multimodality,Learning Representation Learning, Time-Series or Sequence Modeling, Missing data/modality,

Context:

The increasing availability of heterogeneous multimodal data offers new opportunities for artificial intelligence to model complex systems and predict the evolution of their states over time. Such data may originate from multiple sources and exhibit different structures, dimensionalities, sampling rates, and statistical properties.

However, conventional machine learning approaches often rely on observations acquired at a single time point or consider a limited number of modalities. In real-world dynamic systems, observations may instead be collected over time, at irregular intervals, and different modalities may not always be simultaneously available, and some observations may be partially or entirely missing [Huang, 2025] [Lin, 2025]. Moreover, each modality may provide only a partial and complementary view of the underlying state of the observed system.

Recent advances in multimodal learning, temporal modeling, and representation learning provide new opportunities to jointly exploit these heterogeneous observations. Rather than processing each modality or observation independently, these approaches aim to learn unified representations that capture complementary information across modalities as well as relevant dependencies over time.

Recent work on irregular multivariate time series has also highlighted the difficulty of jointly modeling temporal and cross-variable dependencies when observations are unaligned and irregularly sampled [Li, 2025].

In this context, a fundamental challenge is to learn a dynamic latent representation capable of characterizing the current state of a system from heterogeneous and potentially incomplete observations, while preserving the information required to model and predict its future evolution.

Problem statement and aim:

The central challenge of multimodal artificial intelligence is no longer merely to combine multiple data sources, but to learn a reliable representation of a dynamic system from observations that are inherently partial, heterogeneous and collected over time.

Recent approaches to incomplete multimodal learning have investigated different strategies, including missing-modality reconstruction, prompt-based adaptation, multimodal generalization and refinement, and deficiency-resistant representation learning [Lang, 2025] [Huang, 2025] [Lin, 2025]. Despite these advances, learning consistent and reliable representations under varying modality availability remains an open challenge. In particular, approaches based on the reconstruction of missing modalities may face difficulties in preserving the reliability and distributional consistency of reconstructed information [Dai, 2026]. These difficulties become even more important when observations are irregular, asynchronous, noisy, or subject to distribution shifts.

Consequently, an important open research question remains: How can a unified dynamic representation be learned directly from heterogeneous and partially observed data while remaining robust, uncertainty-aware, and interpretable?

The main objective of this thesis is to develop trustworthy multimodal representation learning approaches capable of modeling dynamic states and predicting their future evolution from incomplete temporal observations. Rather than treating missing observations solely as a reconstruction problem, the proposed research will investigate representation-centric learning strategies that preserve complementary information across available modalities while adapting to irregular observations and varying modality availability.

Particular attention will be devoted to four complementary properties of trustworthy prediction: robustness to partial observations and distribution shifts, uncertainty quantification, confidence calibration, and temporal explainability.

Approach:

To achieve the stated objectives, the research strategy will be structured around two main phases.

1- Multimodal and Temporal Representation Learning

The first phase will investigate representation-learning approaches for modeling dynamic states from heterogeneous and partially observed multimodal data. The objective is to learn a unified latent representation that captures complementary information across available modalities while preserving relevant temporal dependencies. Rather than systematically reconstructing missing modalities, the proposed research will investigate representation-centric strategies capable of constructing a reliable latent state directly from the observations that are available. Particular attention will be given to preserving shared and modality-specific information, adapting the learned representation to different modality configurations, and limiting representation shifts caused by missing observations. Recent work on incomplete multimodal learning demonstrates the importance of learning robust and consistent representations when modalities are unavailable [Huang, 2025] [Lin, 2025], while more recent approaches further investigate representation consistency under missing-modality conditions [Chen, 2026].

The temporal dimension will be integrated directly into the representation-learning process. In particular, the research will consider irregular sampling and asynchronous observations, for which temporal observations and modalities may not be aligned. Recent work on irregular multivariate time-series modeling confirms that jointly capturing temporal dependencies and dependencies among unaligned observations remains a significant learning challenge [Li, 2025].

The resulting representation should therefore characterize the current latent state from the available observations while retaining relevant information from previous states. This dynamic representation will then support the prediction of future states under varying data-availability and temporal conditions.

2- Robust, Uncertainty-Aware and Explainable Prediction

The second phase will investigate the trustworthiness of the learned representations and their associated predictions under incomplete, noisy, and shifted data conditions. Robustness will be evaluated through controlled perturbation scenarios involving different patterns and levels of missing modalities, input degradation, and distribution shifts.

Beyond predictive performance, uncertainty estimation will be investigated to determine whether the model can identify situations in which the available information is insufficient or unreliable. Recent approaches have highlighted the relevance of explicitly accounting for uncertainty when learning from incomplete multimodal observations [Nguyen, 2025]. Particular attention will be given to predictive uncertainty and confidence calibration, with the objective of ensuring that model confidence appropriately reflects prediction reliability [Cheon, 2026].

The relationship between robustness and uncertainty will constitute an important component of this investigation. A trustworthy model should maintain stable predictions when sufficient information remains available, while appropriately increasing its uncertainty when degradation, missingness, or distribution shift prevents reliable prediction.

Finally, explainability approaches will be investigated to characterize the contribution of individual modalities and temporal observations to the learned representation and resulting predictions. The objective will be to determine not only which information influences a prediction, but also when this information becomes relevant during the evolution of the observed system.

Together, these components aim to establish a general framework for trustworthy dynamic-state modeling and prediction from incomplete multimodal and temporal observations. The specific tasks of this thesis are:

  • State-of-the-art analysis and problem formulation: Review and analyze recent approaches in multimodal and temporal representation learning, with particular attention to incomplete modalities, irregular and asynchronous observations, and trustworthy prediction. The main limitations of existing approaches will be identified and the research problems addressed in the thesis will be formalized.
  • Dynamic multimodal representation learning: Develop representation-centric approaches for learning unified and dynamic latent representations from heterogeneous and partially observed multimodal data, while preserving complementary information across available modalities and temporal observations.
  • Trustworthy modeling and prediction: Develop and evaluate mechanisms for improving robustness to missing modalities, noise, and distribution shifts, while integrating uncertainty quantification, confidence calibration, and temporal explainability into the proposed models.
  • Experimental evaluation and validation: Evaluate the proposed approaches on multiple multimodal and temporal datasets under different modality-availability, temporal-irregularity, perturbation, and distribution-shift scenarios, considering predictive performance, robustness, uncertainty calibration, explainability, and generalization.

References:

[Lang, 2025] Lang, J., Cheng, Z., Zhong, T., & Zhou, F. (2025). Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 39(17), 18035–18043. DOI: 10.1609/aaai.v39i17.33984.

[Huang, 2025] Huang, W., Chen, Y., Jiang, X., Gao, C., Zhang, T., Chen, Q., & Wang, Y. (2025). Mitigating Pervasive Modality Absence Through Multimodal Generalization and Refinement. Proceedings of the AAAI Conference on Artificial Intelligence, 39(25), 26796–26804. DOI: 10.1609/aaai.v39i25.34883.

[Lin, 2025] Lin, H., Tang, X., Li, H., et al. (2025). T²DR: A Two-Tier Deficiency-Resistant Framework for Incomplete Multimodal Learning. Findings of the Association for Computational Linguistics: ACL 2025, 8602–8616. DOI: 10.18653/v1/2025.findings-acl.452.

[Li, 2025] Li, B., Luo, Y., Liu, Z., Zheng, J., Lv, J., & Ma, Q. (2025). HyperIMTS: Hypergraph Neural Network for Irregular Multivariate Time Series Forecasting. Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), PMLR 267, 35502–35518.

[Nguyen, 2025] Nguyen, D.A., Do, Q.H., Doan, K.D., & Do, M.N. (2025). Are you SURE? Enhancing Multimodal Pretraining with Missing Modalities through Uncertainty Estimation. arXiv:2504.13465.

[Chen, 2026] Chen, J., Cheng, S., Yutao, Y., Zhang, Y., Yuan, H., Peng, P., & Zhong, Y. (2026). PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities. Proceedings of the AAAI Conference on Artificial Intelligence, 40(24), 20076–20082. DOI: 10.1609/aaai.v40i24.39093.

[Dai, 2026] Dai, R., Jian, A., Zhang, R., et al. (2026). Multimodal Learning with Missing Modalities: Can Progressive Diffusion Achieve Distribution-Consistent Learning from Incomplete Multimodal Data? Knowledge-Based Systems, Article 116272. DOI: 10.1016/j.knosys.2026.116272.

[Cheon, 2026] Cheon, J., & Paik, S.-B. (2026). Brain-Inspired Warm-Up Training with Random Noise for Uncertainty Calibration. Nature Machine Intelligence, 8, 602–613.

15/10/2026

  • Diplôme : Master’s degree (research-oriented) or an Engineering degree in computer science
  • Compétences scientifiques : We are looking for a candidate with a strong background in Machine Learning and Deep Learning.The ideal candidate should have:
    • a strong mathematical background, particularly in linear algebra, probability, statistics, and optimization;
    • very good programming skills in Python;
    • the ability to understand, implement, and critically analyze recent machine learning research papers;
    • previous research experience, ideally through a Master’s thesis, research internship, or scientific publication.
  • Experience in at least one of the following areas would be particularly appreciated:
  • Self-Supervised Learning;
  • Time-Series or Sequence Modeling;
  • Probabilistic Machine Learning
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Doctorant en IA - Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction (H/F)
Doctorant en IA - Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction (H/F)

EFREI • Villejuif

Hybride
EUR 29 000 - 36 000
PhD Position F/M Pretrained models of multimodal neuroimaging for predicting individual cognition
PhD Position F/M Pretrained models of multimodal neuroimaging for predicting individual cognition

Inria • Palaiseau

Hybride
EUR 18 000 - 30 000
Remboursement des frais de transport
7 semaines de congés + RTT
Télétravail possible et organisation d
+2
PhD Position F/M Trustable Machine Learning : Analysis and Verification of Soft Automata
PhD Position F/M Trustable Machine Learning : Analysis and Verification of Soft Automata

Inria • Rennes

Sur place
EUR 23 000 - 28 000
Remboursement partiel des transports
Congés annuels + RTT
Télétravail possible après 6 mois
+1
PhD: 3D-Aware Multi-Object Tracking
PhD: 3D-Aware Multi-Object Tracking

IDEMIA Public Security • Courbevoie

Sur place
EUR 55 000 - 85 000
CDD – Post-doctorant IA & Mathématiques Appliquées pour la Prévision Météorologique Long Terme [...]
CDD – Post-doctorant IA & Mathématiques Appliquées pour la Prévision Météorologique Long Terme [...]

TotalEnergies • Palaiseau

Sur place
EUR 42 000 - 66 000
Doctorant F/H Vers des modèles de diffusion efficaces
Doctorant F/H Vers des modèles de diffusion efficaces

Inria • Lyon

Hybride
EUR 20 000 - 27 000
Restauration subventionnée
Transports publics remboursés
Mutuelle et prévoyance
+1
Data Scientist - H/F
Data Scientist - H/F

Mpdata • Balma

Hybride
EUR 45 000 - 75 000
Offre de thèse en TAL/NLP : reconnaissance et liage d’entités pour l’extraction de connaissance[...]
Offre de thèse en TAL/NLP : reconnaissance et liage d’entités pour l’extraction de connaissance[...]

Association française pour l'Intelligence Artificielle • Jouy-en-Josas

Sur place
EUR 28 000 - 35 000
Machine Learning Engineer / Scientist – Search & Retrieval – H/F
Machine Learning Engineer / Scientist – Search & Retrieval – H/F

Anakine.io • Paris

Sur place
EUR 70 000 - 110 000
Prime de vacances
Prévoyance 100%
Remote flexible
+1
Ingénieur·e IA / Machine Learning Confirmé·e
Ingénieur·e IA / Machine Learning Confirmé·e

GoMind • Paris

Sur place
EUR 50 000 - 70 000