TextLens: large language models-powered visual analytics enhancing text clustering
摘要
Text clustering is a cornerstone task in natural language processing with a broad spectrum of applications. Given the advancements in large language models, employing such models to enhance general text clustering has shown promising potential in boosting clustering effectiveness. However, current LLMs-driven approaches often act as black boxes in analyzing the processes of text clustering, leading to poor interpretability. Additionally, these approaches are associated with significant API usage costs and lack effective techniques to explore cluster details. To align these challenges, we propose an LLMs-powered visual analytics approach, called TextLens, to enhance text clustering. First, we present an LLMs-powered framework that integrated LLMs for guiding topic extraction, anomaly filtering, and modification assessment. Second, we introduce a visual analytics system designed to support proposed framework, which facilitates interactive exploration of clusters, analysis of cluster-level thematic extraction, and iterative refinement of clustering results. Finally, we conduct evaluations by applying two datasets into four case studies and a user study to compare clustering outcomes with previous methods, demonstrating the effectiveness and scalability of our approach.
Graphical abstract