<p>Multimodal Sentiment Analysis (MSA) aims to infer human affective states by integrating information from diverse modalities such as text, audio, and vision. Despite recent advances in representation learning and fusion strategies, existing methods often overlook the inherent frequency characteristics within each modality–particularly in audio and visual signals–where task-relevant information may reside in distinct spectral bands. To address this limitation, we propose a novel Frequency-Aware Experts and Multi-Stage Fusion (FEMF) framework. Specifically, we introduce a frequency-aware expert module that decomposes modality-private features into high- and low-frequency components via Discrete Fourier Transform (DFT), and processes them through dedicated expert networks before adaptive fusion. Additionally, we design a multi-stage integration pipeline that incorporates shared-private disentanglement, multi-query modality interaction, and confidence-aware fusion with hierarchical prediction, enabling flexible and robust representation learning across modalities. Extensive experiments on CMU-MOSI and CMU-MOSEI benchmarks demonstrate that our approach achieves superior performance, validating the effectiveness and resilience of the proposed frequency-aware modeling paradigm. Codes are realised at <a href="https://github.com/L11yc/FEMF">https://github.com/L11yc/FEMF</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Frequency-aware experts with multi-stage fusion for multimodal sentiment analysis

  • Xiaofei Zhu,
  • Yaochen Li

摘要

Multimodal Sentiment Analysis (MSA) aims to infer human affective states by integrating information from diverse modalities such as text, audio, and vision. Despite recent advances in representation learning and fusion strategies, existing methods often overlook the inherent frequency characteristics within each modality–particularly in audio and visual signals–where task-relevant information may reside in distinct spectral bands. To address this limitation, we propose a novel Frequency-Aware Experts and Multi-Stage Fusion (FEMF) framework. Specifically, we introduce a frequency-aware expert module that decomposes modality-private features into high- and low-frequency components via Discrete Fourier Transform (DFT), and processes them through dedicated expert networks before adaptive fusion. Additionally, we design a multi-stage integration pipeline that incorporates shared-private disentanglement, multi-query modality interaction, and confidence-aware fusion with hierarchical prediction, enabling flexible and robust representation learning across modalities. Extensive experiments on CMU-MOSI and CMU-MOSEI benchmarks demonstrate that our approach achieves superior performance, validating the effectiveness and resilience of the proposed frequency-aware modeling paradigm. Codes are realised at https://github.com/L11yc/FEMF