Multimodal Sarcasm Detection: A Survey of Methods, Fusion Techniques, Dataset Analysis, and Open Issues
摘要
Sarcasm, expressing negativity through affirmative language, is prevalent in online communication, necessitating accurate detection. Errors in sarcasm detection can distort sentiment analysis, leading to opposite conclusions. While single-modal methods struggle with sarcasm’s subtleties, multimodal techniques offer greater accuracy by capturing inconsistencies across text, environment and facial expressions, despite challenges in integrating modalities. This study examines existing research on sarcasm detection using text-image modalities, discussing available datasets and categorizing models based on fusion techniques and technical frameworks. Some limitations and gaps in existing research are identified based on this review. An important observation is the use of one particular multimodal dataset in most works post-2020, emphasizing the importance of creating diverse, “labeled datasets” for multimodal sarcasm detection. Hybrid fusion with contrastive learning marks a milestone. It boosts MSD performance with the additional advantage of self-learning methodology.