MHAN: Bottleneck Fusion Model Based on Hybrid Attention Network for Multimodal Emotion Recognition
摘要
Multimodal emotion recognition (MER) offers a more comprehensive and accurate understanding of human emotions by analyzing and integrating information from various modalities. However, previous work may have yet to fully consider the emotion feature information within and between modalities, which is the most significant challenge in this domain. Therefore, a novel multimodal hybrid attention network (MHAN) is proposed to achieve more effective information aggregation between text and audio. MHAN includes three main modules, i.e., hybrid encoder block (HEB), multi-head cross-attention (MCA) block, and bottleneck fusion (BF) block. Specifically, the HEB can extract advanced feature representation from text and audio sequences. Moreover, the MCA block can capture common features between text and audio to support emotion recognition. To filter out some redundant information, the BF block is designed to better extract the complementary and most relevant information between text and audio modalities. To validate the effectiveness of MHAN, we have conducted extensive experiments on the IEMOCAP and MELD datasets, and the experimental results demonstrate that MHAN outperforms existing state-of-the-art methods and better interprets subjective human emotions in the applications.