<p>Multimodal Aspect-Based Sentiment Classification (MABSC) aims to analyze the sentiment polarity of aspect terms through text and image modalities. However, the MABSC task still faces challenges such as low-quality modal alignment, insufficient feature fusion, and noise interference. To address these issues, we propose a novel method called Prompt Alignment and Multi-Granularity Feature Fusion (PAMFF). First, to effectively align different modalities, we employ prompt templates to construct prompts for aspect terms and then generate soft entity pseudo-labels derived from both the prompt features and the image entity features. Second, to achieve deeper cross-modal information fusion, we implement fine-grained feature fusion to capture detailed local information and employ a dynamic gating mechanism to adaptively weight each modality based on its contribution. Concurrently, we perform coarse-grained fusion to integrate text-level global semantics with image-level overall scene features. Furthermore, to mitigate noise, we introduce a fusion network layer that strengthens critical sentiment information in the multimodal fused features through dimensional transformation. We also employ a contrastive learning mechanism to cluster semantically consistent cross-modal features in the embedding space, thereby effectively mitigating background noise and irrelevant features. Extensive experiments on three public benchmark datasets demonstrate that the proposed PAMFF outperforms state-of-the-art baselines on MABSC tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PAMFF: Prompt alignment and multi-granularity feature fusion for multimodal aspect-based sentiment classification

  • Yaru Shang,
  • Xuesong Bai,
  • Yimeng Zhan,
  • Donghong Han,
  • Haoyu Yang,
  • Jing Li,
  • Deji Zhao,
  • Gang Wu,
  • Baiyou Qiao

摘要

Multimodal Aspect-Based Sentiment Classification (MABSC) aims to analyze the sentiment polarity of aspect terms through text and image modalities. However, the MABSC task still faces challenges such as low-quality modal alignment, insufficient feature fusion, and noise interference. To address these issues, we propose a novel method called Prompt Alignment and Multi-Granularity Feature Fusion (PAMFF). First, to effectively align different modalities, we employ prompt templates to construct prompts for aspect terms and then generate soft entity pseudo-labels derived from both the prompt features and the image entity features. Second, to achieve deeper cross-modal information fusion, we implement fine-grained feature fusion to capture detailed local information and employ a dynamic gating mechanism to adaptively weight each modality based on its contribution. Concurrently, we perform coarse-grained fusion to integrate text-level global semantics with image-level overall scene features. Furthermore, to mitigate noise, we introduce a fusion network layer that strengthens critical sentiment information in the multimodal fused features through dimensional transformation. We also employ a contrastive learning mechanism to cluster semantically consistent cross-modal features in the embedding space, thereby effectively mitigating background noise and irrelevant features. Extensive experiments on three public benchmark datasets demonstrate that the proposed PAMFF outperforms state-of-the-art baselines on MABSC tasks.