PAMFF: Prompt alignment and multi-granularity feature fusion for multimodal aspect-based sentiment classification
摘要
Multimodal Aspect-Based Sentiment Classification (MABSC) aims to analyze the sentiment polarity of aspect terms through text and image modalities. However, the MABSC task still faces challenges such as low-quality modal alignment, insufficient feature fusion, and noise interference. To address these issues, we propose a novel method called Prompt Alignment and Multi-Granularity Feature Fusion (PAMFF). First, to effectively align different modalities, we employ prompt templates to construct prompts for aspect terms and then generate soft entity pseudo-labels derived from both the prompt features and the image entity features. Second, to achieve deeper cross-modal information fusion, we implement fine-grained feature fusion to capture detailed local information and employ a dynamic gating mechanism to adaptively weight each modality based on its contribution. Concurrently, we perform coarse-grained fusion to integrate text-level global semantics with image-level overall scene features. Furthermore, to mitigate noise, we introduce a fusion network layer that strengthens critical sentiment information in the multimodal fused features through dimensional transformation. We also employ a contrastive learning mechanism to cluster semantically consistent cross-modal features in the embedding space, thereby effectively mitigating background noise and irrelevant features. Extensive experiments on three public benchmark datasets demonstrate that the proposed PAMFF outperforms state-of-the-art baselines on MABSC tasks.