<p>Multimodal emotion recognition in conversation (MERC) involves predicting the emotion category of a conversation on the basis of textual, acoustic, and visual modalities. Information from these diverse modalities can reinforce each other to enhance the accuracy of emotion prediction. However, some information modalities may be absent in real-world applications and information from various modalities may be difficult to integrate. Therefore, a suitable strategy is required to compensate for missing modalities by using information from the available modalities and prioritizing important information. Consequently, this study developed a graph convolutional network (GCN) model with a feature compensation module and dual-channel second-order pooling module for MERC. This model initially uses a GCN to compensate for missing features by aggregating features corresponding to the same utterance node. Subsequently, it applies dual-channel second-order pooling to sift through and integrate all features. Empirical evaluations of the proposed model against other baseline models on two benchmark datasets, namely the IEMOCAP and MELD datasets, indicated that the proposed model outperformed the other models; this finding underscores the effectiveness of the proposed model in recognizing the emotions expressed in multimodal dialogue data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph convolutional network model with a feature compensation module and dual-channel second-order pooling module for multimodal emotion recognition in conversation

  • Xiaocong Tan,
  • Zhengze Gong,
  • Mengkun Gan,
  • Weijie Xie,
  • Wenhui Wang

摘要

Multimodal emotion recognition in conversation (MERC) involves predicting the emotion category of a conversation on the basis of textual, acoustic, and visual modalities. Information from these diverse modalities can reinforce each other to enhance the accuracy of emotion prediction. However, some information modalities may be absent in real-world applications and information from various modalities may be difficult to integrate. Therefore, a suitable strategy is required to compensate for missing modalities by using information from the available modalities and prioritizing important information. Consequently, this study developed a graph convolutional network (GCN) model with a feature compensation module and dual-channel second-order pooling module for MERC. This model initially uses a GCN to compensate for missing features by aggregating features corresponding to the same utterance node. Subsequently, it applies dual-channel second-order pooling to sift through and integrate all features. Empirical evaluations of the proposed model against other baseline models on two benchmark datasets, namely the IEMOCAP and MELD datasets, indicated that the proposed model outperformed the other models; this finding underscores the effectiveness of the proposed model in recognizing the emotions expressed in multimodal dialogue data.