Strategies for Addressing the Limited Labeled Datasets in Fake News Detection: A Systematic Review
摘要
The demand for effective fake news detection methods is increasing, necessitating strategies to address the constraints of human-labeled datasets. This review aims to determine several techniques for handling unlabeled data, evaluate the performance of state-of-the-art methods, and identify future research recommendations. Three hundred sixty articles were collected, and 64 were selected from four databases using the Kitchenham protocol and the Parsifal application. The review’s findings figure out suitable methods and techniques considering the source, type, and quantity of data. The current literature highlights the commonly used techniques, including intrinsically semi-supervised, wrapper, graph-based, and hybrid methods that combine transductive and inductive approaches. The most advanced methods exhibited robust performance, with several achieving F1 scores of 0.9 or above. Several prospective study ideas have been identified, such as improving the algorithm and data handling scenarios, with recent advancements in Large Language Models (LLMs) offering the transformative potential to address these areas through enhanced feature extraction, synthetic data generation, and cross-lingual generalization. The significance of these findings lies in their ability to effectively summarize and reveal strategies and prospects for detecting fake news with limited labeled data.