AHSD: adaptive multimodal fusion with hierarchical semantic disentanglement for sentiment analysis
摘要
Multimodal Sentiment Analysis (MSA) faces two persistent challenges: the presence of deceptive or noisy information within modalities, and the representational bottlenecks of shallow prediction heads. To address these challenges, we propose an Adaptive Multimodal Fusion with Hierarchical Semantic Disentanglement (AHSD), which achieves robust multimodal representation learning through uncertainty-aware dynamic calibration, alongside semantic disentanglement and refinement. Specifically, to mitigate the impact of deceptive or noisy information within modalities, the Adaptive Multimodal Attention Mechanism (AMMA) performs uncertainty-aware modality calibration by learning sample-specific reliability-aware weights for each modality, thereby enabling adaptive and robust cross-modal interaction. To overcome the limitations of shallow prediction, the Hierarchical Feature Refinement (HFR) module adopts a semantic disentanglement and refinement strategy based on parallel subspace projection. HFR maps the fused representations into complementary semantic subspaces. This enables the model to simultaneously capture coarse-grained sentiment semantics and fine-grained affective cues without substantially increasing model complexity. Furthermore, we employ a dual-regularization strategy (attention entropy and modality balance) to actively penalize single-modality dominance and ensure stable optimization. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that AHSD consistently achieves highly competitive performance against state-of-the-art multimodal sentiment analysis approaches across multiple evaluation metrics.