Aspect-Based Multimodal Sentiment Analysis (ABMSA) is a fine-grained mission that intents to analyze users’ sentiment expression towards target aspects through different and rich modal contents. In conjunction with this task, many methods have been proposed to link modalities to form interactive judgments of emotional tendencies. However, the proposed method still has some limitations: (1) Since different modalities exist in different feature spaces, the differences between modalities are large and it is difficult to align them; (2) It is difficult for modalities to interact during the fusion process. To resolve these points, we propose a novel ABMSA network model that achieves alignment between modalities by sharing prompt parameters and improves the fusion performance to achieve effective classification through hierarchical fusion. Specifically, to solve the alignment problem between different modalities, we propose shared source prompt parameters to fine-tune the model branches to achieve modality alignment. To improve the fusion performance between modalities, aspect-aware attention is used in the lower stages and fusion layers are used in the higher levels to achieve high fusion between modalities. The outcomes of our experiments indicate that the model we developed attains leading-edge performance levels when tested on three datasets. Furthermore, a large number of experiments have shown that the model we put forward exhibits remarkable performance and robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Modal Shared Prompts for Aspect-Based Multimodal Sentiment Analysis

  • Jiachang Sun,
  • Jiaxuan Sun

摘要

Aspect-Based Multimodal Sentiment Analysis (ABMSA) is a fine-grained mission that intents to analyze users’ sentiment expression towards target aspects through different and rich modal contents. In conjunction with this task, many methods have been proposed to link modalities to form interactive judgments of emotional tendencies. However, the proposed method still has some limitations: (1) Since different modalities exist in different feature spaces, the differences between modalities are large and it is difficult to align them; (2) It is difficult for modalities to interact during the fusion process. To resolve these points, we propose a novel ABMSA network model that achieves alignment between modalities by sharing prompt parameters and improves the fusion performance to achieve effective classification through hierarchical fusion. Specifically, to solve the alignment problem between different modalities, we propose shared source prompt parameters to fine-tune the model branches to achieve modality alignment. To improve the fusion performance between modalities, aspect-aware attention is used in the lower stages and fusion layers are used in the higher levels to achieve high fusion between modalities. The outcomes of our experiments indicate that the model we developed attains leading-edge performance levels when tested on three datasets. Furthermore, a large number of experiments have shown that the model we put forward exhibits remarkable performance and robustness.