<p>Transformer-based Large Language Models (LLMs), which have recently gained popularity, have significantly impacted the various fields of natural language processing. One of them is text clustering, which involves categorizing the huge texts produced by today’s digital world into meaningful groups. LLMs enable text clustering with a more semantic and contextualized approach than traditional methods. One such model is Sentence-BERT (SBERT), which has been modified to detect semantic similarity between texts. Before being clustered, texts need to be transformed into numerical text embeddings. SBERT-based models have shown promise in generating meaningful sentence embeddings. However, they face limitations when dealing with long texts that exceed their maximum token limit. In this context, this study proposes two distinct methods to overcome these limitations and enhance the performance of SBERT models for clustering long text. The proposed methods are combined with various SBERT models, and their combinations are compared to the existing default method on three datasets containing lengthy texts. This study evaluates the impact of these methods on the models and their contributions to clustering performance. The findings indicate that the proposed methods exhibit a higher clustering performance of up to 14% than the default method in text clustering. Additionally, this study provides valuable insights into the text clustering performance of SBERT models, offering practical implications for further research and applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing SBERT for long text clustering: two novel approaches with empirical insights

  • Yasin Ortakci,
  • Burak Borhan

摘要

Transformer-based Large Language Models (LLMs), which have recently gained popularity, have significantly impacted the various fields of natural language processing. One of them is text clustering, which involves categorizing the huge texts produced by today’s digital world into meaningful groups. LLMs enable text clustering with a more semantic and contextualized approach than traditional methods. One such model is Sentence-BERT (SBERT), which has been modified to detect semantic similarity between texts. Before being clustered, texts need to be transformed into numerical text embeddings. SBERT-based models have shown promise in generating meaningful sentence embeddings. However, they face limitations when dealing with long texts that exceed their maximum token limit. In this context, this study proposes two distinct methods to overcome these limitations and enhance the performance of SBERT models for clustering long text. The proposed methods are combined with various SBERT models, and their combinations are compared to the existing default method on three datasets containing lengthy texts. This study evaluates the impact of these methods on the models and their contributions to clustering performance. The findings indicate that the proposed methods exhibit a higher clustering performance of up to 14% than the default method in text clustering. Additionally, this study provides valuable insights into the text clustering performance of SBERT models, offering practical implications for further research and applications.