As misinformation continues to spread rapidly on social media platforms identifying and stopping the dissemination of fake news has become an urgent need. In this article, we propose a deep learning approach leveraging keywords for feature extraction and classification of Arabic dialect fake news. Our method achieves an accuracy of 82.3% on a corpus comprising 3000 news articles in Algerian and Tunisian dialects, Modern Standard Arabic (MSA), French, and English, featuring instances of code-switching between these languages; as well as an accuracy of 93.7% on an English fake news corpus. Our experimentation shows that the shortcut learning problem that can arise when using keyword based features can be solved using regularization techniques. Our findings also show that our approach will achieve better performance on larger Arabic dialect corpora.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Fake News: Exploring Key Features in Multilingual Arabic Dialect Corpus

  • Abdelouahab Hocini,
  • Kamel Smaili

摘要

As misinformation continues to spread rapidly on social media platforms identifying and stopping the dissemination of fake news has become an urgent need. In this article, we propose a deep learning approach leveraging keywords for feature extraction and classification of Arabic dialect fake news. Our method achieves an accuracy of 82.3% on a corpus comprising 3000 news articles in Algerian and Tunisian dialects, Modern Standard Arabic (MSA), French, and English, featuring instances of code-switching between these languages; as well as an accuracy of 93.7% on an English fake news corpus. Our experimentation shows that the shortcut learning problem that can arise when using keyword based features can be solved using regularization techniques. Our findings also show that our approach will achieve better performance on larger Arabic dialect corpora.