We are in the information age, but also in the era of disinformation, with millions of fake news items circulating daily. Various fields are working to identify and understand fake news. We focus on hybrid approaches combining machine learning and natural language processing, using surface linguistic features, which are independent of language and enable a multilingual approach. Many studies rely on binary classification, overlooking multiclass problems and class imbalance, often focusing only on English. We propose a methodology that applies surface linguistic features for multiclass fake news detection in a multilingual context. Experiments were conducted on two datasets, LIAR (English) and CLNews (Spanish), both imbalanced. Using Synthetic Minority Oversampling Technique (SMOTE), Random Oversampling (ROS), and Random Undersampling (RUS), we observed improved class detection. For example, in LIAR, the classification of the “false” class improved by 43.38% using SMOTE with Adaptive Boosting. In CLNews, the ROS technique with Random Forest raised accuracy to 95%, representing a 158% relative improvement over the unbalanced scenario. These results highlight our approach’s effectiveness in addressing the problem of multiclass fake news detection in an imbalanced, multilingual context.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Surface Linguistic Features for Multiclass Fake News Detection in a Multilingual Context

  • Eduardo Puraivan,
  • Pablo Ormeño-Arriagada,
  • Steffanie Kloss,
  • Connie Cofré-Morales

摘要

We are in the information age, but also in the era of disinformation, with millions of fake news items circulating daily. Various fields are working to identify and understand fake news. We focus on hybrid approaches combining machine learning and natural language processing, using surface linguistic features, which are independent of language and enable a multilingual approach. Many studies rely on binary classification, overlooking multiclass problems and class imbalance, often focusing only on English. We propose a methodology that applies surface linguistic features for multiclass fake news detection in a multilingual context. Experiments were conducted on two datasets, LIAR (English) and CLNews (Spanish), both imbalanced. Using Synthetic Minority Oversampling Technique (SMOTE), Random Oversampling (ROS), and Random Undersampling (RUS), we observed improved class detection. For example, in LIAR, the classification of the “false” class improved by 43.38% using SMOTE with Adaptive Boosting. In CLNews, the ROS technique with Random Forest raised accuracy to 95%, representing a 158% relative improvement over the unbalanced scenario. These results highlight our approach’s effectiveness in addressing the problem of multiclass fake news detection in an imbalanced, multilingual context.