<p>Detecting knowledge emerging trends has received increasing attention. It can help researchers understand the history of the discipline and predict future research hotspots. Dynamic topic models can be used to identify knowledge emerging trends from academic papers. However, traditional dynamic topic models have some shortcomings, such as over-assumptions, insufficient topic distinction, and high computational cost. To address this problem, we propose a relevance-based dynamic thin topic model (RBDTTM). We model topic evolution with a Gaussian process and adopt a relevance-based mechanism on topic-word distributions. Under this assumption, only words relevant to a certain topic can be represented. This relevance-based mechanism can not only decrease the number of parameters to be estimated but also achieve more prominent and focused topics. We evaluate the estimation performance of RBDTTM using a series of experiments on synthetic data. Results show that RBDTTM has greater interpretability and generalization than its competitors. Finally, we take the statistics discipline as an example and apply RBDTTM to two corpora of journal articles and a Chinese graduation thesis to explore the emerging statistical knowledge trend in the past two decades.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Knowledge emerging trend detection using relevance-based dynamic thin topic model

  • Feifei Wang,
  • Xueqiong Yuan,
  • Xiaoling Lu

摘要

Detecting knowledge emerging trends has received increasing attention. It can help researchers understand the history of the discipline and predict future research hotspots. Dynamic topic models can be used to identify knowledge emerging trends from academic papers. However, traditional dynamic topic models have some shortcomings, such as over-assumptions, insufficient topic distinction, and high computational cost. To address this problem, we propose a relevance-based dynamic thin topic model (RBDTTM). We model topic evolution with a Gaussian process and adopt a relevance-based mechanism on topic-word distributions. Under this assumption, only words relevant to a certain topic can be represented. This relevance-based mechanism can not only decrease the number of parameters to be estimated but also achieve more prominent and focused topics. We evaluate the estimation performance of RBDTTM using a series of experiments on synthetic data. Results show that RBDTTM has greater interpretability and generalization than its competitors. Finally, we take the statistics discipline as an example and apply RBDTTM to two corpora of journal articles and a Chinese graduation thesis to explore the emerging statistical knowledge trend in the past two decades.