Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.
Association Bernard Gregory propose un doctorat sur le sujet: Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction. Le contrat doctoral porte sur l’apprentissage multimodal et les modèles dynamiques à partir de données incomplètes, dans un cadre de Villejuif.
Le doctorant travaillera sur des représentations fiables et adaptatives selon les modalités disponibles. Le projet vise à développer des représentations latentes robustes, gérer l’irrégularité des
19/08/2026 Contrat doctoral
Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction
Multimodality,Learning Representation Learning, Time-Series or Sequence Modeling, Missing data/modality,
Context:
The increasing availability of heterogeneous multimodal data offers new opportunities for artificial intelligence to model complex systems and predict the evolution of their states over time. Such data may originate from multiple sources and exhibit different structures, dimensionalities, sampling rates, and statistical properties.
However, conventional machine learning approaches often rely on observations acquired at a single time point or consider a limited number of modalities. In real-world dynamic systems, observations may instead be collected over time, at irregular intervals, and different modalities may not always be simultaneously available, and some observations may be partially or entirely missing [Huang, 2025] [Lin, 2025]. Moreover, each modality may provide only a partial and complementary view of the underlying state of the observed system.
Recent advances in multimodal learning, temporal modeling, and representation learning provide new opportunities to jointly exploit these heterogeneous observations. Rather than processing each modality or observation independently, these approaches aim to learn unified representations that capture complementary information across modalities as well as relevant dependencies over time.
Recent work on irregular multivariate time series has also highlighted the difficulty of jointly modeling temporal and cross-variable dependencies when observations are unaligned and irregularly sampled [Li, 2025].
In this context, a fundamental challenge is to learn a dynamic latent representation capable of characterizing the current state of a system from heterogeneous and potentially incomplete observations, while preserving the information required to model and predict its future evolution.
Problem statement and aim:
The central challenge of multimodal artificial intelligence is no longer merely to combine multiple data sources, but to learn a reliable representation of a dynamic system from observations that are inherently partial, heterogeneous and collected over time.
Recent approaches to incomplete multimodal learning have investigated different strategies, including missing-modality reconstruction, prompt-based adaptation, multimodal generalization and refinement, and deficiency-resistant representation learning [Lang, 2025] [Huang, 2025] [Lin, 2025]. Despite these advances, learning consistent and reliable representations under varying modality availability remains an open challenge. In particular, approaches based on the reconstruction of missing modalities may face difficulties in preserving the reliability and distributional consistency of reconstructed information [Dai, 2026]. These difficulties become even more important when observations are irregular, asynchronous, noisy, or subject to distribution shifts.
Consequently, an important open research question remains: How can a unified dynamic representation be learned directly from heterogeneous and partially observed data while remaining robust, uncertainty-aware, and interpretable?
The main objective of this thesis is to develop trustworthy multimodal representation learning approaches capable of modeling dynamic states and predicting their future evolution from incomplete temporal observations. Rather than treating missing observations solely as a reconstruction problem, the proposed research will investigate representation-centric learning strategies that preserve complementary information across available modalities while adapting to irregular observations and varying modality availability.
Particular attention will be devoted to four complementary properties of trustworthy prediction: robustness to partial observations and distribution shifts, uncertainty quantification, confidence calibration, and temporal explainability.
Approach:
To achieve the stated objectives, the research strategy will be structured around two main phases.
1- Multimodal and Temporal Representation Learning
The first phase will investigate representation-learning approaches for modeling dynamic states from heterogeneous and partially observed multimodal data. The objective is to learn a unified latent representation that captures complementary information across available modalities while preserving relevant temporal dependencies. Rather than systematically reconstructing missing modalities, the proposed research will investigate representation-centric strategies capable of constructing a reliable latent state directly from the observations that are available. Particular attention will be given to preserving shared and modality-specific information, adapting the learned representation to different modality configurations, and limiting representation shifts caused by missing observations. Recent work on incomplete multimodal learning demonstrates the importance of learning robust and consistent representations when modalities are unavailable [Huang, 2025] [Lin, 2025], while more recent approaches further investigate representation consistency under missing-modality conditions [Chen, 2026].
The temporal dimension will be integrated directly into the representation-learning process. In particular, the research will consider irregular sampling and asynchronous observations, for which temporal observations and modalities may not be aligned. Recent work on irregular multivariate time-series modeling confirms that jointly capturing temporal dependencies and dependencies among unaligned observations remains a significant learning challenge [Li, 2025].
The resulting representation should therefore characterize the current latent state from the available observations while retaining relevant information from previous states. This dynamic representation will then support the prediction of future states under varying data-availability and temporal conditions.
2- Robust, Uncertainty-Aware and Explainable Prediction
The second phase will investigate the trustworthiness of the learned representations and their associated predictions under incomplete, noisy, and shifted data conditions. Robustness will be evaluated through controlled perturbation scenarios involving different patterns and levels of missing modalities, input degradation, and distribution shifts.
Beyond predictive performance, uncertainty estimation will be investigated to determine whether the model can identify situations in which the available information is insufficient or unreliable. Recent approaches have highlighted the relevance of explicitly accounting for uncertainty when learning from incomplete multimodal observations [Nguyen, 2025]. Particular attention will be given to predictive uncertainty and confidence calibration, with the objective of ensuring that model confidence appropriately reflects prediction reliability [Cheon, 2026].
The relationship between robustness and uncertainty will constitute an important component of this investigation. A trustworthy model should maintain stable predictions when sufficient information remains available, while appropriately increasing its uncertainty when degradation, missingness, or distribution shift prevents reliable prediction.
Finally, explainability approaches will be investigated to characterize the contribution of individual modalities and temporal observations to the learned representation and resulting predictions. The objective will be to determine not only which information influences a prediction, but also when this information becomes relevant during the evolution of the observed system.
Together, these components aim to establish a general framework for trustworthy dynamic-state modeling and prediction from incomplete multimodal and temporal observations. The specific tasks of this thesis are:
References:
[Lang, 2025] Lang, J., Cheng, Z., Zhong, T., & Zhou, F. (2025). Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 39(17), 18035–18043. DOI: 10.1609/aaai.v39i17.33984.
[Huang, 2025] Huang, W., Chen, Y., Jiang, X., Gao, C., Zhang, T., Chen, Q., & Wang, Y. (2025). Mitigating Pervasive Modality Absence Through Multimodal Generalization and Refinement. Proceedings of the AAAI Conference on Artificial Intelligence, 39(25), 26796–26804. DOI: 10.1609/aaai.v39i25.34883.
[Lin, 2025] Lin, H., Tang, X., Li, H., et al. (2025). T²DR: A Two-Tier Deficiency-Resistant Framework for Incomplete Multimodal Learning. Findings of the Association for Computational Linguistics: ACL 2025, 8602–8616. DOI: 10.18653/v1/2025.findings-acl.452.
[Li, 2025] Li, B., Luo, Y., Liu, Z., Zheng, J., Lv, J., & Ma, Q. (2025). HyperIMTS: Hypergraph Neural Network for Irregular Multivariate Time Series Forecasting. Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), PMLR 267, 35502–35518.
[Nguyen, 2025] Nguyen, D.A., Do, Q.H., Doan, K.D., & Do, M.N. (2025). Are you SURE? Enhancing Multimodal Pretraining with Missing Modalities through Uncertainty Estimation. arXiv:2504.13465.
[Chen, 2026] Chen, J., Cheng, S., Yutao, Y., Zhang, Y., Yuan, H., Peng, P., & Zhong, Y. (2026). PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities. Proceedings of the AAAI Conference on Artificial Intelligence, 40(24), 20076–20082. DOI: 10.1609/aaai.v40i24.39093.
[Dai, 2026] Dai, R., Jian, A., Zhang, R., et al. (2026). Multimodal Learning with Missing Modalities: Can Progressive Diffusion Achieve Distribution-Consistent Learning from Incomplete Multimodal Data? Knowledge-Based Systems, Article 116272. DOI: 10.1016/j.knosys.2026.116272.
[Cheon, 2026] Cheon, J., & Paik, S.-B. (2026). Brain-Inspired Warm-Up Training with Random Noise for Uncertainty Calibration. Nature Machine Intelligence, 8, 602–613.
15/10/2026