The Effect of Different Diacritic Settings on the Accuracy of Sentiment Analysis in Czech and Slovak Slavic Languages Upon Transfer-Learning
摘要
Sentiment analysis is still actively developed scientific field of research. Vast majority of the publications (dedicated to Czech or Slovak languages) focus on the evaluation of models during training phase and perform no transfer-learning evaluations. Diacritical aspects of Slavic languages are also omitted during data preprocessing. In this paper we focus on the transfer-learning evaluation, evaluating a pre-trained sentiment analysis model in another domain. For this purpose, we selected two popular pre-trained state-of-the-art deep-learning sentiment analysis models Czert and SlovakBERT along with our traditional machine learning Maximum Entropy model. Upon evaluation we put attention on different diacritical settings of the datasets. We notice that the current two deep-learning models actually differ in the reaction to different diacritic settings. We did find robustness but also slight variations in F scores. We theorize that more attention should be paid in diacritical setting in datasets used to train models for sentiment analysis. Also, more transfer-learning evaluations of the sentiment analysis models should be performed in order to test for the robustness of the models.