The rise of fake news on digital platforms has led to a growing interest in the development of automatic detection models. This research evaluates the effectiveness of different natural language processing (NLP) algorithms to identify fake news in Spanish. For this purpose, we implemented five encoding models: GloVe, BART, XLNet, DeBerta, and RoBERTa, combined with six classifier models: LSTM, GRU, DNC, CNN, BETO and ANN. Additionally, we develop a proprietary Spanish-language dataset which includes real and fake news collected from digital newspapers and social media. The models were evaluated using accuracy, recall, F1-score, and AUC-ROC metrics. The GloVe-DNC model achieved the best overall performance, with an accuracy of 0.96, an F1-score of 0.97, a recall of 0.98, and an AUC-ROC of 0.98. In contrast, the BART-LSTM and RoBERTa-ANN models showed the lowest overall results, with AUC-ROC scores of 0.63 and 0.71, respectively. These findings contribute to the field of fake news detection in Spanish and highlight potential directions for further research, such as expanding the dataset and exploring new model architectures to enhance detection accuracy and effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fake News Detection in a Real-World Spanish Dataset: A Neural Network and Transformer-Based Approach

  • Fiorella Nina,
  • Angelina Arana,
  • Edwin Escobedo,
  • Guillermo Dávila

摘要

The rise of fake news on digital platforms has led to a growing interest in the development of automatic detection models. This research evaluates the effectiveness of different natural language processing (NLP) algorithms to identify fake news in Spanish. For this purpose, we implemented five encoding models: GloVe, BART, XLNet, DeBerta, and RoBERTa, combined with six classifier models: LSTM, GRU, DNC, CNN, BETO and ANN. Additionally, we develop a proprietary Spanish-language dataset which includes real and fake news collected from digital newspapers and social media. The models were evaluated using accuracy, recall, F1-score, and AUC-ROC metrics. The GloVe-DNC model achieved the best overall performance, with an accuracy of 0.96, an F1-score of 0.97, a recall of 0.98, and an AUC-ROC of 0.98. In contrast, the BART-LSTM and RoBERTa-ANN models showed the lowest overall results, with AUC-ROC scores of 0.63 and 0.71, respectively. These findings contribute to the field of fake news detection in Spanish and highlight potential directions for further research, such as expanding the dataset and exploring new model architectures to enhance detection accuracy and effectiveness.