This study explores the day-to-day activities of cryptocurrency (Bitcoin) traders who might have gotten their crowd knowledge of trade from news articles, thereby shaping the trend of the market prices. The latent Dirichlet allocation (LDA) model and its variant, the supervised latent Dirichlet allocation model (sLDA), were used to analyze 4073 preprocessed, scraped news articles from CNBC’s market section website and compared them for prediction purposes. The models were trained and tested using the document-term matrix alongside “k” various values 3,10,20,30,50,100, and 200. Due to our classification method of multinomial classification, we employed the services of four metric evaluators, namely, mean absolute percentage error (MAPE), mean absolute error (MAE), root mean square error (RMSE), and the R2. The result shows that the sLDA model will outperform the LDA model + (classification or regression model) in terms of the prediction task of labels for unlabeled new documents’ purpose considering the metric evaluation values.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparison Between the Extrapolation Strengths of Unsupervised and Supervised Topic Models

  • T. O. Maku,
  • M. O. Adenomon,
  • M. U. Adehi

摘要

This study explores the day-to-day activities of cryptocurrency (Bitcoin) traders who might have gotten their crowd knowledge of trade from news articles, thereby shaping the trend of the market prices. The latent Dirichlet allocation (LDA) model and its variant, the supervised latent Dirichlet allocation model (sLDA), were used to analyze 4073 preprocessed, scraped news articles from CNBC’s market section website and compared them for prediction purposes. The models were trained and tested using the document-term matrix alongside “k” various values 3,10,20,30,50,100, and 200. Due to our classification method of multinomial classification, we employed the services of four metric evaluators, namely, mean absolute percentage error (MAPE), mean absolute error (MAE), root mean square error (RMSE), and the R2. The result shows that the sLDA model will outperform the LDA model + (classification or regression model) in terms of the prediction task of labels for unlabeled new documents’ purpose considering the metric evaluation values.