<p>The spread of hate speech has become a serious issue in online communities, negatively influencing user interaction and platform integrity. Although major social media companies have devoted substantial effort to developing automated detection systems, reliably classifying hateful content continues to be difficult. One key reason is shortcut learning, where models rely on frequent or sensitive words instead of understanding sentence meaning and context. Many hate speech detection systems achieve high accuracy while relying on surface-level lexical cues rather than robust semantic understanding. This shortcut-based behavior can lead to unstable predictions, particularly for neutral texts containing identity-related terms. In this paper, we analyze the prediction stability of transformer-based hate speech models when classifying neutral content. We construct a bilingual dataset of English and Norwegian social media posts annotated for binary hate speech detection and evaluate several Norwegian-specific and multilingual transformer models under identical training conditions. To mitigate shortcut learning, we introduce a neutral only stability regularization objective that encourages consistent predictions under stochastic perturbations. In post-training, we assess residual instability by applying controlled masking-based input variations through random token masking and measuring prediction variability. Our results show that Norwegian-specific models, especially Nor-BERT<InlineEquation ID="IEq1"><EquationSource Format="TEX">\(_{\text {base}}\)</EquationSource></InlineEquation>, produce more stable predictions for neutral content while maintaining competitive classification performance. In contrast, the best-performing multilingual model, mm-BERT<InlineEquation ID="IEq2"><EquationSource Format="TEX">\(_{\text {base}}\)</EquationSource></InlineEquation>, achieves the highest overall accuracy of 86%, outperforming the Norwegian-specific models in aggregate classification accuracy. However, all models exhibit high worst-case instability, indicating that shortcut-based behavior is not fully resolved. These findings demonstrate that standard accuracy metrics alone are insufficient and highlight the importance of stability-aware training and evaluation when deploying hate speech detection systems in real-world moderation tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exposing shortcut sensitivity in English–Norwegian hate speech detection via neutral stability analysis

  • Ehtesham Hashmi,
  • Muhammad Tayyab Mazhar,
  • Sule Yildirim Yayilgan,
  • Rajendra Akerkar,
  • Mehtab Afzal

摘要

The spread of hate speech has become a serious issue in online communities, negatively influencing user interaction and platform integrity. Although major social media companies have devoted substantial effort to developing automated detection systems, reliably classifying hateful content continues to be difficult. One key reason is shortcut learning, where models rely on frequent or sensitive words instead of understanding sentence meaning and context. Many hate speech detection systems achieve high accuracy while relying on surface-level lexical cues rather than robust semantic understanding. This shortcut-based behavior can lead to unstable predictions, particularly for neutral texts containing identity-related terms. In this paper, we analyze the prediction stability of transformer-based hate speech models when classifying neutral content. We construct a bilingual dataset of English and Norwegian social media posts annotated for binary hate speech detection and evaluate several Norwegian-specific and multilingual transformer models under identical training conditions. To mitigate shortcut learning, we introduce a neutral only stability regularization objective that encourages consistent predictions under stochastic perturbations. In post-training, we assess residual instability by applying controlled masking-based input variations through random token masking and measuring prediction variability. Our results show that Norwegian-specific models, especially Nor-BERT\(_{\text {base}}\), produce more stable predictions for neutral content while maintaining competitive classification performance. In contrast, the best-performing multilingual model, mm-BERT\(_{\text {base}}\), achieves the highest overall accuracy of 86%, outperforming the Norwegian-specific models in aggregate classification accuracy. However, all models exhibit high worst-case instability, indicating that shortcut-based behavior is not fully resolved. These findings demonstrate that standard accuracy metrics alone are insufficient and highlight the importance of stability-aware training and evaluation when deploying hate speech detection systems in real-world moderation tasks.