Text, vision and audio are three typical modalities for multimodal sentiment analysis. The work BERT is an often-used backbone to fuse text with others. However, approaches using BERT all more or less ignore some information like outputs of BERT’s middle layers and unimodal information discarded because of fusion. We believe that both unimodal and multimodal features have their own unique information beneficial for sentiment judgment. In this paper, we propose a new framework, GFoR, which fuse representations greedily to retain as much useful information as possible. To leverage the rich semantic information in BERT’s middle layers without disturbing the unimodal encoding, we design the parallel fusion for multimodal interaction. In this way, the purity of unimodal representations and the sufficient fusion for BERT intermediate layers can be simultaneously ensured. Finally, we fuse all the unimodal and bimodal representations for prediction. In addition to direct integration into multimodal representations, external implicit fusion, a cross-modal generation task, is used to suppress irrelevant and conflicting information from visual and acoustic modalities. The experimental results on CMU-MOSI and CMU-MOSEI demonstrate the effectiveness of our approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Greedy Fusion Oriented Representations for Multimodal Sentiment Analysis

  • Ran Xu

摘要

Text, vision and audio are three typical modalities for multimodal sentiment analysis. The work BERT is an often-used backbone to fuse text with others. However, approaches using BERT all more or less ignore some information like outputs of BERT’s middle layers and unimodal information discarded because of fusion. We believe that both unimodal and multimodal features have their own unique information beneficial for sentiment judgment. In this paper, we propose a new framework, GFoR, which fuse representations greedily to retain as much useful information as possible. To leverage the rich semantic information in BERT’s middle layers without disturbing the unimodal encoding, we design the parallel fusion for multimodal interaction. In this way, the purity of unimodal representations and the sufficient fusion for BERT intermediate layers can be simultaneously ensured. Finally, we fuse all the unimodal and bimodal representations for prediction. In addition to direct integration into multimodal representations, external implicit fusion, a cross-modal generation task, is used to suppress irrelevant and conflicting information from visual and acoustic modalities. The experimental results on CMU-MOSI and CMU-MOSEI demonstrate the effectiveness of our approach.