Hypergraph-based multimodal adaptive fusion for emotion recognition in conversation
摘要
Emotion recognition in conversation (ERC) aims to identify and understand emotional expressions in conversation, which plays an important role in improving the human–computer interaction experience. How to effectively model both the speaker and conversational context remains a significant challenge in ERC. Existing approaches has primarily focused on graph-based methods to address this issue. Graph neural network (GNN) has shown remarkable effectiveness in capturing relational information within data. However, they typically capture only pairwise relationships, restricting GNN-based methods in handling higher-order semantics and complex multimodal interactions within conversational contexts. In this paper, we propose a hypergraph-based multimodal adaptive fusion method for ERC. We construct three intra-modal hypergraphs for textual modality, visual modality, and acoustic modality in conversational utterances, and introduce contextual hyperedges and speaker self-dependency hyperedges into them. From the two perspectives of conversational context utterances dependence and speaker emotional dependence, the model can more comprehensively understand the emotional information in the conversation. Meanwhile, we construct an inter-modal hypergraph to capture the interactions between different modalities. Additionally, considering the differences in representation capabilities of different modalities, we employ an attention-based gated neural network to fuse the hypergraph information after hypergraph aggregation, in order to reduce redundancy and noise during the cross-modal fusion process. Experimental results on two public datasets verify the superiority of the proposed approach over state-of-the-art approaches.