Robust multimodal emotion recognition under missing and incomplete data with cross-modal regeneration
-
Aims: Multimodal emotion recognition (MER) can outperform unimodal approaches by integrating complementary information from multiple sources. However, real-world applications often involve incomplete or missing modalities, reducing the reliability ...
MoreAims: Multimodal emotion recognition (MER) can outperform unimodal approaches by integrating complementary information from multiple sources. However, real-world applications often involve incomplete or missing modalities, reducing the reliability of existing MER models. This study proposes a framework that remains robust under missing-modality conditions while preserving the advantages of multimodal integration.
Methods: To address this challenge, we propose a cross-modal latent regeneration and attention long short-term memory (CMLR-ALSTM) framework. The framework combines pretrained variational autoencoder encoders with residual projection networks, optimized with L2 loss, to achieve stable latent-space alignment across modalities. Regenerated and available latent representations are integrated using a cross-modal attention mechanism and processed by an LSTM to model temporal dependencies and improve multimodal fusion under incomplete data.
Results: The proposed framework was evaluated on three benchmark datasets under complete, partially missing, and completely missing modality scenarios. Experimental results demonstrate that CMLR-ALSTM achieves up to 17.22% improvement under missing-modality conditions while maintaining competitive performance under complete-modality settings. The results also show that the proposed latent regeneration strategy effectively preserves cross-modal relationships and robust latent representations.
Conclusion: The experimental results confirm the effectiveness of the proposed framework, particularly in realistic environments where data availability is inconsistent. By leveraging CMLR to reconstruct missing representations and modelling temporal dependencies through LSTM, the proposed approach provides a more robust and reliable MER framework for practical deployment. The results across different modality combinations demonstrate its ability to generalize across heterogeneous multimodal settings.
Less -
Behzad Mahaseni, Naimul Mefraz Khan
-
DOI: https://doi.org/10.70401/ec.2026.0023 - July 02, 2026
-
This article belongs to the Special Issue Adaptive Empathic Interactive Media for Therapy