In Tunisia, citizens use social media platforms as a space to exercise freedom of speech. However, unchecked and complete freedom of expression can fuel the spread of hateful speech, which is devastating not only for those targeted but also for our society. This alarming situation evokes the need for limiting the spread of hateful content by working on hate speech detection in “Derja”, which is the tunisian dialect. Used as a means of communication in daily life and on social media platforms, this dialect is a mixture of many languages, including Arabic, French, and Amazighi, and it can be written using Arabic letters. Due to the complexity of this language, a significant lack of publicly available, large, and annotated datasets for hate speech detection in Tunisian dialect written in Arabic letters is noticeable, making “Tunisian Derja” an underrepresented dialect. In this paper, we introduce the largest publicly available dataset, which consists of more than 12k comments manually annotated as Hate, and Neutral. We also provide an in-depth explanation of the processes of data collection, annotation, and pre-processing. Moreover, we undertake a comprehensive evaluation of the dataset’s efficacy through various machine learning models, including Support Vector Machines (SVM), Random Forest, and XGBoost.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HateTune: Tunisian Dialect Hate Speech Detection Dataset

  • Ons Kharrat,
  • Fatma Alzahra Mohamed,
  • Ikram Mtimet,
  • Nour Benamor,
  • Chayma Fourati

摘要

In Tunisia, citizens use social media platforms as a space to exercise freedom of speech. However, unchecked and complete freedom of expression can fuel the spread of hateful speech, which is devastating not only for those targeted but also for our society. This alarming situation evokes the need for limiting the spread of hateful content by working on hate speech detection in “Derja”, which is the tunisian dialect. Used as a means of communication in daily life and on social media platforms, this dialect is a mixture of many languages, including Arabic, French, and Amazighi, and it can be written using Arabic letters. Due to the complexity of this language, a significant lack of publicly available, large, and annotated datasets for hate speech detection in Tunisian dialect written in Arabic letters is noticeable, making “Tunisian Derja” an underrepresented dialect. In this paper, we introduce the largest publicly available dataset, which consists of more than 12k comments manually annotated as Hate, and Neutral. We also provide an in-depth explanation of the processes of data collection, annotation, and pre-processing. Moreover, we undertake a comprehensive evaluation of the dataset’s efficacy through various machine learning models, including Support Vector Machines (SVM), Random Forest, and XGBoost.