Graph convolutional network model with a feature compensation module and dual-channel second-order pooling module for multimodal emotion recognition in conversation
摘要
Multimodal emotion recognition in conversation (MERC) involves predicting the emotion category of a conversation on the basis of textual, acoustic, and visual modalities. Information from these diverse modalities can reinforce each other to enhance the accuracy of emotion prediction. However, some information modalities may be absent in real-world applications and information from various modalities may be difficult to integrate. Therefore, a suitable strategy is required to compensate for missing modalities by using information from the available modalities and prioritizing important information. Consequently, this study developed a graph convolutional network (GCN) model with a feature compensation module and dual-channel second-order pooling module for MERC. This model initially uses a GCN to compensate for missing features by aggregating features corresponding to the same utterance node. Subsequently, it applies dual-channel second-order pooling to sift through and integrate all features. Empirical evaluations of the proposed model against other baseline models on two benchmark datasets, namely the IEMOCAP and MELD datasets, indicated that the proposed model outperformed the other models; this finding underscores the effectiveness of the proposed model in recognizing the emotions expressed in multimodal dialogue data.