<p>Text clustering is a cornerstone task in natural language processing with a broad spectrum of applications. Given the advancements in large language models, employing such models to enhance general text clustering has shown promising potential in boosting clustering effectiveness. However, current LLMs-driven approaches often act as black boxes in analyzing the processes of text clustering, leading to poor interpretability. Additionally, these approaches are associated with significant API usage costs and lack effective techniques to explore cluster details. To align these challenges, we propose an LLMs-powered visual analytics approach, called TextLens, to enhance text clustering. First, we present an LLMs-powered framework that integrated LLMs for guiding topic extraction, anomaly filtering, and modification assessment. Second, we introduce a visual analytics system designed to support proposed framework, which facilitates interactive exploration of clusters, analysis of cluster-level thematic extraction, and iterative refinement of clustering results. Finally, we conduct evaluations by applying two datasets into four case studies and a user study to compare clustering outcomes with previous methods, demonstrating the effectiveness and scalability of our approach.</p> Graphical abstract

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TextLens: large language models-powered visual analytics enhancing text clustering

  • Ruixiao Peng,
  • Yu Dong,
  • Guan Li,
  • Dong Tian,
  • Guihua Shan

摘要

Text clustering is a cornerstone task in natural language processing with a broad spectrum of applications. Given the advancements in large language models, employing such models to enhance general text clustering has shown promising potential in boosting clustering effectiveness. However, current LLMs-driven approaches often act as black boxes in analyzing the processes of text clustering, leading to poor interpretability. Additionally, these approaches are associated with significant API usage costs and lack effective techniques to explore cluster details. To align these challenges, we propose an LLMs-powered visual analytics approach, called TextLens, to enhance text clustering. First, we present an LLMs-powered framework that integrated LLMs for guiding topic extraction, anomaly filtering, and modification assessment. Second, we introduce a visual analytics system designed to support proposed framework, which facilitates interactive exploration of clusters, analysis of cluster-level thematic extraction, and iterative refinement of clustering results. Finally, we conduct evaluations by applying two datasets into four case studies and a user study to compare clustering outcomes with previous methods, demonstrating the effectiveness and scalability of our approach.

Graphical abstract