Topic modeling offers a useful way to examine the topic labels of extensive document collections, facilitating the organization and outline of the themes within that collection. Previous researchers have suggested considering the probabilistic model, where each document is the convex combination of topic vectors, and the topic vector is a distribution of words. However, finding an appropriate distribution vector for each topic is not easy for a high-dimensional word co-occurrence space. This work provides an alternative topic vector inference method combined with non-negative matrix factorization for learning high-quality topics. To verify the effectiveness and priority of the proposed method, we experiment with three public benchmark datasets, NIPS, Movies, and NYtimes, and show a competitive performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Anchor Words Inference for Stochastic Matrix Factorization

  • Lingping Kong,
  • Zdenek Dostal,
  • Millie Pant,
  • Václav Snášel

摘要

Topic modeling offers a useful way to examine the topic labels of extensive document collections, facilitating the organization and outline of the themes within that collection. Previous researchers have suggested considering the probabilistic model, where each document is the convex combination of topic vectors, and the topic vector is a distribution of words. However, finding an appropriate distribution vector for each topic is not easy for a high-dimensional word co-occurrence space. This work provides an alternative topic vector inference method combined with non-negative matrix factorization for learning high-quality topics. To verify the effectiveness and priority of the proposed method, we experiment with three public benchmark datasets, NIPS, Movies, and NYtimes, and show a competitive performance.