<p>The task of Multi-modal Multi-label Emotion Recognition (MMER) is to identify various emotions from multiple heterogeneous data sources, including text, audio, and visual data. Existing methods often overlook the contributions of different modality features within a single sample to the final emotion recognition results, as well as the complexity of mixed emotions. This paper introduces a novel framework called ConFusion that utilizes the Shapley value to determine the marginal contribution of each modality at the sample level, and subsequently employs a state space model for multi-modal feature fusion to enhance the model’s ability to focus on the most informative cues. To further address the complexity of emotional expressions, we introduce an emotion decoding mechanism to capture the dependencies among different emotion labels and obtain representative label embeddings. Our method not only enhances the understanding of modality contributions but also provides a fine-grained emotion representation. Experiments conducted on both aligned and unaligned settings of the benchmark MMER dataset, CMU-MOSEI, demonstrate the precision and efficiency of our approach compared with the previous state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Modality Contribution-based Fusion for Multi-modal Multi-label Emotion Recognition

  • Jia Nie,
  • Wenlong Dong,
  • Qirong Mao

摘要

The task of Multi-modal Multi-label Emotion Recognition (MMER) is to identify various emotions from multiple heterogeneous data sources, including text, audio, and visual data. Existing methods often overlook the contributions of different modality features within a single sample to the final emotion recognition results, as well as the complexity of mixed emotions. This paper introduces a novel framework called ConFusion that utilizes the Shapley value to determine the marginal contribution of each modality at the sample level, and subsequently employs a state space model for multi-modal feature fusion to enhance the model’s ability to focus on the most informative cues. To further address the complexity of emotional expressions, we introduce an emotion decoding mechanism to capture the dependencies among different emotion labels and obtain representative label embeddings. Our method not only enhances the understanding of modality contributions but also provides a fine-grained emotion representation. Experiments conducted on both aligned and unaligned settings of the benchmark MMER dataset, CMU-MOSEI, demonstrate the precision and efficiency of our approach compared with the previous state-of-the-art methods.