A new fuzzy document clustering algorithm based on topic homogeneity is introduced. In detail, a novel dissimilarity measure is proposed, derived from the p-value of a hypothesis test that assesses the homogeneity of topic distributions between two documents. First, the topic distributions are derived through Latent Dirichlet Allocation, and then a bootstrap procedure is applied to obtain the p-value. Finally, the resulting dissimilarity matrix is integrated into the fuzzy relational clustering procedure. The performance of the proposal is evaluated using a benchmark dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topic Homogeneity Test-Based Fuzzy Document Clustering

  • Gian Mario Sangiovanni,
  • Louisa Kontoghiorghes,
  • Ana Colubi,
  • Maria Brigida Ferraro

摘要

A new fuzzy document clustering algorithm based on topic homogeneity is introduced. In detail, a novel dissimilarity measure is proposed, derived from the p-value of a hypothesis test that assesses the homogeneity of topic distributions between two documents. First, the topic distributions are derived through Latent Dirichlet Allocation, and then a bootstrap procedure is applied to obtain the p-value. Finally, the resulting dissimilarity matrix is integrated into the fuzzy relational clustering procedure. The performance of the proposal is evaluated using a benchmark dataset.