<p>Topic modeling remains a key approach for uncovering latent thematic structures in large text corpora. Classical Bag-of-Words (BoW) models such as LDA and NMF now coexist with embedding-based methods like Top2Vec and BERTopic, prompting renewed assessment of their respective strengths and limitations. This study advances this comparison in two ways. First, building on a previously proposed BoW-based method combining Growing Neural Gas clustering and feature maximization (CFMf), we introduce a reclassification variant (CFMf-R) to refine ill-formed clusters, an embedding-based version (CFMf-e-R) that integrates contextual text representations and a dimensionality-reduced variant (CFMf-e-U-R) that applies UMAP to the contextual embeddings. Second, using a corpus of 16,917 philosophy of science research articles, we develop a pipeline to systematically evaluate eight topic modeling approaches—LDA, NMF, CFMf-R, CFMf-e-R, CFMf-e-U-R, Top2Vec, BERTopic, and HD-BERTopic—across coherence, diversity, and recall metrics, while also examining document distributions and topic interpretability. Results reveal distinct trade-offs: Top2Vec achieves high coherence and diversity but low recall and interpretability; BERTopic slightly outperforms LDA in coherence but not in recall. Document distributions appear more regular with LDA, CFMf-e-R, and CFMf-e-U-R; CFMf-e-R and BERTopic yield higher interpretability. Overall, CFMf-e-R tends to offer a better balance across dimensions, though no single model dominates. Notably, topic content and size vary substantially across methods, revealing distinct corpus perspectives. Reaffirming the continuing relevance of BoW-based models and emphasizing the modularity of topic modeling pipelines, these studies suggest that ensemble approaches combining multiple models may yield more robust corpus representations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bag-of-words or embedding-based topic models? Comparing LDA, BERTopic, and six other methods

  • Jean-Charles Lamirel,
  • Francis Lareau,
  • Thibault Prouteau,
  • Christophe Malaterre

摘要

Topic modeling remains a key approach for uncovering latent thematic structures in large text corpora. Classical Bag-of-Words (BoW) models such as LDA and NMF now coexist with embedding-based methods like Top2Vec and BERTopic, prompting renewed assessment of their respective strengths and limitations. This study advances this comparison in two ways. First, building on a previously proposed BoW-based method combining Growing Neural Gas clustering and feature maximization (CFMf), we introduce a reclassification variant (CFMf-R) to refine ill-formed clusters, an embedding-based version (CFMf-e-R) that integrates contextual text representations and a dimensionality-reduced variant (CFMf-e-U-R) that applies UMAP to the contextual embeddings. Second, using a corpus of 16,917 philosophy of science research articles, we develop a pipeline to systematically evaluate eight topic modeling approaches—LDA, NMF, CFMf-R, CFMf-e-R, CFMf-e-U-R, Top2Vec, BERTopic, and HD-BERTopic—across coherence, diversity, and recall metrics, while also examining document distributions and topic interpretability. Results reveal distinct trade-offs: Top2Vec achieves high coherence and diversity but low recall and interpretability; BERTopic slightly outperforms LDA in coherence but not in recall. Document distributions appear more regular with LDA, CFMf-e-R, and CFMf-e-U-R; CFMf-e-R and BERTopic yield higher interpretability. Overall, CFMf-e-R tends to offer a better balance across dimensions, though no single model dominates. Notably, topic content and size vary substantially across methods, revealing distinct corpus perspectives. Reaffirming the continuing relevance of BoW-based models and emphasizing the modularity of topic modeling pipelines, these studies suggest that ensemble approaches combining multiple models may yield more robust corpus representations.