Text Enhancement-Based Multimodal Fusion for Video Sentiment Analysis
摘要
Learning multimodal features of videos is a key for accurately understanding the real emotion expressed in the videos. Sequence alignment is usually necessary to deal with different modal sequence lengths in order to fuse multimodal features together. Text Enhancement-based Multimodal Fusion (TEMF) is implemented in the paper to integrate multimodal information without sequence alignment for improving video sentiment analysis. For modeling the unaligned sequences of multimodal inputs, text is enhanced by cross-modal attention mechanism. The expression level of non-textual features is strengthened so that textual and non-textual features interact at similar representational level. Experiments are conducted with three benchmark datasets and demonstrate TEMF outperforms several comparison methods in terms of Acc and F1.