A cross-validation study of Turkish sentiment analysis datasets and tools
摘要
In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability and reuse of Turkish datasets across studies has yielded highly diverse outcomes. To address this, we conducted a systematic review of sentiment analysis studies on Turkish text. Our search identified 78 relevant studies, from which we extracted over 80 datasets. These studies were labeled using a comprehensive sentiment analysis taxonomy, and the dataset details were compiled into a structured repository. Furthermore, we evaluated the performance of four state-of-the-art models-XLM-T, BERTurk (fine-tuned with the BounTi dataset), TSAM, and TurkishBERTweet-on four widely-used Turkish datasets. Among the models, XLM-T achieved the highest performance with an accuracy of 0.92 and F1 score of 0.95 on the Twt dataset, while TSAM reached 0.97 accuracy and F1 score on the Humir dataset. Our empirical results demonstrate that model performance varies significantly based on dataset characteristics such as domain, balance, and linguistic structure. Our review revealed key research gaps, including the limited application of emotion-based and concept-based sentiment analysis techniques and the lack of domain diversity in Turkish sentiment datasets. By highlighting such gaps and compiling a centralized repository, this study provides a comprehensive and publicly accessible resource to guide future research in Turkish sentiment analysis.