A Diffusion Driven Multimodal Fusion Framework for Context Aware Sarcasm Detection via Sentiment Syntax Graph Modeling
摘要
Sarcasm, a complex form of figurative expression, often conveys meanings that deviate from its literal interpretation. Detecting sarcasm in online communication presents unique challenges due to the intricate relationships between textual and visual elements. Existing multimodal sarcasm detection models primarily focus on localized inconsistencies but fail to capture deeper contextual and semantic dependencies between text and images. To address these limitations, this study introduces the multimodal sarcasm-aware fusion network (MSAFN), an advanced framework that integrates multiple novel components for enhanced sarcasm recognition. The proposed model incorporates a context-aware multimodal alignment module (CMAM), which dynamically refines text-image feature interactions to improve sarcasm-related pattern recognition. Additionally, a graph-driven sentiment discrepancy module (GSDM) leverages graph-based sentiment analysis and syntactic dependency structures to enhance sentiment-text coherence, improving the identification of multimodal inconsistencies. Using the SenticNet knowledge base gives the model prior emotional understanding, helping it align text and image information more effectively. A Unified Cross-Attention Transformer (X-Former) facilitates efficient multimodal fusion by aligning textual and visual features at both global and local levels. Furthermore, an image captioning mechanism is integrated to reinforce the alignment of semantic and emotional cues between modalities, strengthening the contextual coherence of sarcasm interpretation. Experimental evaluations on benchmark datasets show that MSAFN consistently outperforms previous state-of-the-art models. By effectively integrating global multimodal relationships with fine-grained sentiment-inconsistency modeling, the proposed framework enhances cross-modal sarcasm prediction robustness. These findings highlight the potential of multimodal fusion strategies in advancing sentiment analysis, particularly in social media monitoring, automated content moderation, and sentiment-aware AI systems.