A progressive interaction model for multimodal sarcasm detection
摘要
Multimodal sarcasm detection aims to determine whether conflicting semantics arise in different modalities. Existing research, primarily relying on direct interaction between image and text, limits model performance in sarcasm detection due to the difficulty in cross-modal alignment and information integration caused by semantic differences between the modalities. In this paper, we propose a progressive interaction approach. First, unlike the traditional direct interaction approach, the pre-interaction approach is adopted by bridging the image and text through attributes to reduce the semantic difference between them. Then, contrastive learning is employed to align image and text features for better synchronization of image-text semantics. Finally, sarcasm cues are captured through the interaction between image and text for detecting sarcasm. In the pre-interaction phase, we design different components for image and text respectively for their interaction with attributes. Experiments demonstrate the excellent performance of our method on a multimodal sarcasm detection task.