Enhancing Segment-Based Bag of Clusters with Mixture Models and Late Chunking for Document Clustering
摘要
Document clustering remains a fundamental task in information retrieval, yet accurately capturing semantic structure in long and context-rich texts poses persistent challenges. In this paper, we propose SMoC-LC (Segment-based Mixture of Clusters with Late Chunking), a novel clustering framework that addresses two key limitations of prior methods: fixed-length segmentation and hard cluster assignments. Our approach introduces Late Chunking to produce flexible, variable-length text segments using long-context embeddings, and employs Gaussian Mixture Models (GMM) to enable soft-probabilistic clustering. We benchmark SMoC-LC and its variants (SBoC, SBoC-LC, SMoC) on seven datasets spanning different domains and structural complexity, including AGNews, 20News-10K, BBCNews, Reuters-21578, and DBpedia (L1-L3). Results show that SMoC-LC consistently improves clustering quality across accuracy (ACC), normalized mutual information (NMI), and adjusted Rand index (ARI), with statistically significant gains observed in complex, hierarchical datasets. Our analysis reveals that Late Chunking is especially beneficial for short, structured documents, while soft clustering excels in ambiguous or multi-topic contexts. These findings underscore the need for adaptable clustering strategies aligned with textual granularity and semantic ambiguity.