Moroccan Arabic (MA) dialect is a low resource language. To perform any NLP task, we have to develop the necessary resources from scratch. This paper introduces our work on MAOffens, the first MA dataset for offensive language detection. The dataset will serve to build predictive models to detect offensive content widely present on social media and hence help ensure online safety. We built the dataset with a mixture of comments in Arabic and Latin scripts to cover offensiveness in both cases. The resulting dataset consists of 23k comments totally balanced. The dataset is open to the public ( https://huggingface.co/datasets/randa/maoffens ). We evaluated the annotation and classification power of the dataset through various classifier architectures. Our best performing classifier was based on a MA transformer model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MAOffens: Moroccan Arabic Offensive Language Dataset

  • Randa Zarnoufi,
  • Mohammed Hajhouj,
  • Walid Bachri,
  • Hamid Jaafar,
  • Mounia Abik

摘要

Moroccan Arabic (MA) dialect is a low resource language. To perform any NLP task, we have to develop the necessary resources from scratch. This paper introduces our work on MAOffens, the first MA dataset for offensive language detection. The dataset will serve to build predictive models to detect offensive content widely present on social media and hence help ensure online safety. We built the dataset with a mixture of comments in Arabic and Latin scripts to cover offensiveness in both cases. The resulting dataset consists of 23k comments totally balanced. The dataset is open to the public ( https://huggingface.co/datasets/randa/maoffens ). We evaluated the annotation and classification power of the dataset through various classifier architectures. Our best performing classifier was based on a MA transformer model.