Transformer-based Multimodal Sentiment Analysis (MSA) has garnered growing attention for its strong ability to model cross-modal interaction and capture global features. Nevertheless, these methods generate substantial computational and GPU memory consumption, leading to model inefficiency and parameter redundancy. Therefore, we propose an efficient Query-Shared Multimodal Transformer (QSMT) for improving model efficiency and modeling robust cross-modal interaction. Specifically, we first propose a Query-Shared Cross-Attention (QSCA) mechanism, the core of QSMT, to fuse all modalities into multimodal fusion representation for interaction with unimodal representations, reducing computational complexity and space requirements of previous cross-modal interactions. Then, we further strengthen the supervisory role of the unimodal label generation module, which considers modality-specific information, on cross-modal interaction by applying average pooling. Extensive experimental results on commonly used benchmark datasets, including CMU-MOSI and CMU-MOSEI, demonstrate that the proposed QSMT achieves superior performance and significant efficiency gains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

QSMT: Query-Shared Multimodal Transformer for Multimodal Sentiment Analysis

  • Huazhong Liu,
  • Yuanhan Liu,
  • Jihong Ding,
  • Wenxuan Zhang

摘要

Transformer-based Multimodal Sentiment Analysis (MSA) has garnered growing attention for its strong ability to model cross-modal interaction and capture global features. Nevertheless, these methods generate substantial computational and GPU memory consumption, leading to model inefficiency and parameter redundancy. Therefore, we propose an efficient Query-Shared Multimodal Transformer (QSMT) for improving model efficiency and modeling robust cross-modal interaction. Specifically, we first propose a Query-Shared Cross-Attention (QSCA) mechanism, the core of QSMT, to fuse all modalities into multimodal fusion representation for interaction with unimodal representations, reducing computational complexity and space requirements of previous cross-modal interactions. Then, we further strengthen the supervisory role of the unimodal label generation module, which considers modality-specific information, on cross-modal interaction by applying average pooling. Extensive experimental results on commonly used benchmark datasets, including CMU-MOSI and CMU-MOSEI, demonstrate that the proposed QSMT achieves superior performance and significant efficiency gains.