Mongolian Multimodal Sentiment Analysis Based on Multi-level Attention and Convolution-Enhanced Fusion
摘要
To address the challenges of insufficient feature extraction, inadequate cross-modal interactions, and multimodal feature fusion in Mongolian Multimodal Sentiment Analysis (MSA), this paper proposes a Mongolian MSA model based on Multi-Level Attention and Convolution-Enhanced Fusion (MACF). First, a Multi-Level Attention progressive mechanism is introduced, which comprises a hierarchical interaction framework consisting of multi-head Self-Attention (SA), Cross-Modal Attention (CMA), and Soft-Attention (Soft-Attn) to enhance feature extraction and strengthen cross-modal interactions. Then, a Convolution-Enhanced Fusion (CEF) module is designed to improve local temporal correlation and suppress cross-modal noise. By complementing the global feature selection of the Soft-Attn mechanism with local feature enhancement, this module enables deep multimodal feature fusion. Experimental results demonstrate that, compared to various state-of-the-art models, the proposed model significantly improves the accuracy of Mongolian multimodal sentiment analysis.