Multimodal sarcasm detection based on cross-modal semantic understanding and knowledge enhancement
摘要
Multimodal sarcasm detection represents a cutting-edge research area at the intersection of natural language processing and computer vision, aiming to integrate textual and visual information to identify sarcastic relationships between verbal expressions and underlying intentions. However, existing research continues to face significant challenges in cross-modal semantic alignment, complex sarcastic context understanding, and knowledge-enhanced utilization, resulting in limited model capabilities for capturing sarcastic semantics. To address these limitations, this paper proposes CMSKE (Cross-modal semantic understanding and knowledge enhancement), a novel multimodal sarcasm detection model that integrates cross-modal semantic understanding with knowledge enhancement mechanisms. The model incorporates a multi-perspective knowledge enhancement strategy, strengthens semantic alignment with a cross-modal attention mechanism, and optimizes the semantic feature space through category center-based contrastive learning. CMSKE comprises 203.83M parameters in total. This paper also analyzes the model’s efficiency performance across different computational environments, validating its scalability on supercomputing platforms. Experimental results on two publicly available datasets, Twitter Sarcasm and MVSA-Multiple, demonstrate that CMSKE exhibits good performance across key evaluation metrics, verifying the effectiveness of the proposed model on the selected datasets.