HateLens: Cognition-inspired contextual reasoning with dynamic cross-modal collaboration for hateful meme detection
摘要
The rapid proliferation of online social platforms has led to the widespread dissemination of hateful memes, which not only pollute the online information ecosystem, but also pose severe challenges to intelligent information management and social governance. Despite the continuous advancement of hateful meme detection techniques, existing models still suffer from two critical limitations: contextual cognition deficit in hate understanding and imbalanced modality collaboration for hate inference, which hinder their practical application in real-world intelligent information systems. To address these issues, we propose a cognition-inspired detection framework, HateLens, tailored for intelligent information processing scenarios, which consists of a knowledge-enhanced contextual reasoning module and a dynamic cross-modal collaboration module. The contextual reasoning module integrates domain-specific knowledge related to hateful content through corpus-driven retrieval, augmented with large language models-guided semantic expansion to enhance the understanding of complex contextual information. The cross-modal collaboration module adopts an adaptive soft voting strategy to dynamically balance the contributions of textual cues, visual evidence, and LLM-augmented hateful clues, and leverages multi-dimensional cognitive alignment to calibrate the weights of different modalities, ensuring effective fusion of multimodal information. Experimental results on two widely used benchmark datasets demonstrate that our proposed HateLens framework outperforms state-of-the-art models in terms of detection accuracy, precision, recall, and F1-score, verifying its effectiveness and applicability in intelligent information systems for hateful meme detection.