<p>Recent research in multimodal sentiment analysis (MSA) has primarily focused on the fusion of modal characteristics and the interactions between multiple modalities. Most existing modal feature interaction methods either directly concatenate features from individual modalities or employ transformer-based approaches to extract and facilitate interaction between them. However, direct fusion methods often fail to capture the interaction between modalities and overlook the central role of text information in sentiment analysis. While transformer models can effectively enable intermodal interaction, they typically result in an increase in model parameters, which can hinder practical deployment. Additionally, noise unrelated to sentiment in non-linguistic modalities can compromise model accuracy. To address these challenges, this paper proposes a multimodal sentiment analysis model based on the All-MLP architecture, called CU-SEMLP. First, we introduce a novel modal interaction module, Shift-MLP, designed to facilitate the sharing of text-based information. Shift-MLP enhances modal interaction through spatial shift operation, and can capture rich emotional information in text-based mixed modalities, improving the expressiveness and adaptability of the model. Second, we propose the EM-MLP module as a replacement for transformer-based approaches, targeting noise suppression in non-linguistic modalities. EM-MLP simulates attention mechanisms through shift operations and a series of linear layers and normalization. Since both modules are based on the All-MLP architecture, the overall model parameter count is significantly reduced. We evaluate CU-SEMLP on the CMU-MOSI and CMU-MOSEI datasets. Compared to baseline methods, CU-SEMLP achieves better performance with fewer parameters, and a series of ablation experiments demonstrate the effectiveness of each module. The experimental results show that CU-SEMLP can effectively complete the multimodal sentiment analysis task.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CU-SEMLP: All-MLP-based multimodal interaction model for multimodal sentiment analysis

  • Siyuan Liu,
  • Hongkun Zhao,
  • Yang Chen,
  • Fanmin Kong,
  • Kang Li

摘要

Recent research in multimodal sentiment analysis (MSA) has primarily focused on the fusion of modal characteristics and the interactions between multiple modalities. Most existing modal feature interaction methods either directly concatenate features from individual modalities or employ transformer-based approaches to extract and facilitate interaction between them. However, direct fusion methods often fail to capture the interaction between modalities and overlook the central role of text information in sentiment analysis. While transformer models can effectively enable intermodal interaction, they typically result in an increase in model parameters, which can hinder practical deployment. Additionally, noise unrelated to sentiment in non-linguistic modalities can compromise model accuracy. To address these challenges, this paper proposes a multimodal sentiment analysis model based on the All-MLP architecture, called CU-SEMLP. First, we introduce a novel modal interaction module, Shift-MLP, designed to facilitate the sharing of text-based information. Shift-MLP enhances modal interaction through spatial shift operation, and can capture rich emotional information in text-based mixed modalities, improving the expressiveness and adaptability of the model. Second, we propose the EM-MLP module as a replacement for transformer-based approaches, targeting noise suppression in non-linguistic modalities. EM-MLP simulates attention mechanisms through shift operations and a series of linear layers and normalization. Since both modules are based on the All-MLP architecture, the overall model parameter count is significantly reduced. We evaluate CU-SEMLP on the CMU-MOSI and CMU-MOSEI datasets. Compared to baseline methods, CU-SEMLP achieves better performance with fewer parameters, and a series of ablation experiments demonstrate the effectiveness of each module. The experimental results show that CU-SEMLP can effectively complete the multimodal sentiment analysis task.