<p>Harmful content on social media, especially in the form of memes, poses unique challenges for moderation and analysis. Memes combine text and imagery to convey complex messages rapidly and evoke emotional responses, enabling the dissemination of harmful ideas, reinforcement of stereotypes, and normalization of discriminatory behaviour, often eluding standard moderation tools. Existing datasets, such as Hateful Memes (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\approx\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≈</mo> </math></EquationSource> </InlineEquation>10 K samples) and Memotion (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\approx\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≈</mo> </math></EquationSource> </InlineEquation>8 K samples) focus on binary or coarse-grained labels and omit many common forms of harm. To address this gap, we introduce GuardHarMem, a new corpus of <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\approx\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≈</mo> </math></EquationSource> </InlineEquation>16.6 K memes annotated with fine-grained categories including racism, mockery, and promotion of harmful substances. We also present HarMDetect a practical baseline multimodal classifier that integrates text, image, and automatically extracted captions. By applying targeted data augmentation strategies, we enhance model robustness on GuardHarMem, HarMDetect outperforms baseline transformers for both binary and multiclass classification. The dataset and code are publicly available at <a href="https://github.com/EL-Amrany/Harmful-memes-detection">https://github.com/EL-Amrany/Harmful-memes-detection</a>. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GuardHarMem and HarMDetect: a multimodal dataset and benchmlark model for fine-grained harmful meme classification

  • Samir El-amrany,
  • Salima Lamsiyah,
  • Matthias R. Brust,
  • Pascal Bouvry

摘要

Harmful content on social media, especially in the form of memes, poses unique challenges for moderation and analysis. Memes combine text and imagery to convey complex messages rapidly and evoke emotional responses, enabling the dissemination of harmful ideas, reinforcement of stereotypes, and normalization of discriminatory behaviour, often eluding standard moderation tools. Existing datasets, such as Hateful Memes ( \(\approx\) 10 K samples) and Memotion ( \(\approx\) 8 K samples) focus on binary or coarse-grained labels and omit many common forms of harm. To address this gap, we introduce GuardHarMem, a new corpus of \(\approx\) 16.6 K memes annotated with fine-grained categories including racism, mockery, and promotion of harmful substances. We also present HarMDetect a practical baseline multimodal classifier that integrates text, image, and automatically extracted captions. By applying targeted data augmentation strategies, we enhance model robustness on GuardHarMem, HarMDetect outperforms baseline transformers for both binary and multiclass classification. The dataset and code are publicly available at https://github.com/EL-Amrany/Harmful-memes-detection.