MuLAD: Multimodal Aggression Detection from Social Media Memes Exploiting Visual and Textual Features
摘要
Aggression detection from memes is challenging due to their region-specific interpretation and multimodal nature. Detecting or classifying aggressive memes is complicated in low-resource languages (including Bengali) because benchmark datasets and primary language processing software are needed. This paper proposes an innovative meme classification technique that harnesses deep learning (DL) approaches to leverage memes’ visual and textual features in Bengali. Various DL frameworks, such as VGG16, VGG19, ResNet50, CNN, BiLSTM, and BiLSTM+CNN, extract visual and textual features from memes. A novel corpus named the Bengali Meme Dataset (AMemD) is also introduced, comprising a substantial amount of multimodal data, including text and image components. Experimental results on AMemD demonstrate the effectiveness of the proposed approach. The CNN combined with VGG16 obtained the highest \(f_1\) -score of 0.738 among all multimodal techniques tested. This pioneering research offers valuable insights into the complex task of aggression detection from memes in Bengali and provides a foundation for future studies in this area. The dataset is available at https://github.com/Maruf089/Multimodal-Aggression-Detection .