Transparent Hate Speech Detection in Norwegian Using Explainable AI
摘要
Social networks are central to communication in the contemporary world. However, the anonymity afforded to users has emboldened the spread of online hate speech (HS) and poses a significant challenge. Online hate speech detection remains a formidable challenge in the digital age, particularly for languages with fewer computational resources like Norwegian. This paper presents a comprehensive study on HS detection utilizing transformer-based models designed for the Norwegian language, enhanced by hyperparameter tuning and regularization. Our investigation not only benchmarks the effectiveness of various models, with Nor-BERT_base emerging as notably superior for smaller datasets but also integrates explainable AI techniques to unveil the model’s decisions using LIME. This approach enhances the transparency of AI-driven content moderation and paves the way for future enhancements involving multilingual capabilities and advanced interpretative models.