<p>The utilization of large language models (LLMs) in research is becoming increasingly prevalent, as they offer advanced capabilities in processing and generating human-like text. However, this advancement comes with a significant trade-off in terms of time and computational costs. In this paper, we demonstrate that analyzing large text datasets with the use of LLMs can be performed efficiently in terms of both time and energy. For this purpose, we utilize the Llama pre-trained model. In more detail, we study the topic modeling task where the goal is to discover and identify topics in large text corpora. The basis of our approach is a hierarchical divisive clustering technique that clusters the data based on their semantic similarity, after employing a Sentence-BERT encoder, pre-trained on a variety of data across different tasks. Then, using an LLM, we identify topics for representative samples from each cluster. Additionally, we introduce a new evaluation method that leverages the capabilities of LLMs to assess the alignment between discovered topics and ground truth labels, providing a robust validation metric. Our findings indicate that it is possible to effectively reduce the computational cost of the topic modeling process compared to the direct application of LLMs and BERTopic, while simultaneously enhancing inference time and overall efficiency, thereby surpassing the current state-of-the-art capabilities of BERTopic.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large language models for efficient topic modeling

  • Panagiotis C. Theocharopoulos,
  • Panagiotis Anagnostou,
  • Spiros V. Georgakopoulos,
  • Sotiris K. Tasoulis,
  • Vassilis P. Plagianakos

摘要

The utilization of large language models (LLMs) in research is becoming increasingly prevalent, as they offer advanced capabilities in processing and generating human-like text. However, this advancement comes with a significant trade-off in terms of time and computational costs. In this paper, we demonstrate that analyzing large text datasets with the use of LLMs can be performed efficiently in terms of both time and energy. For this purpose, we utilize the Llama pre-trained model. In more detail, we study the topic modeling task where the goal is to discover and identify topics in large text corpora. The basis of our approach is a hierarchical divisive clustering technique that clusters the data based on their semantic similarity, after employing a Sentence-BERT encoder, pre-trained on a variety of data across different tasks. Then, using an LLM, we identify topics for representative samples from each cluster. Additionally, we introduce a new evaluation method that leverages the capabilities of LLMs to assess the alignment between discovered topics and ground truth labels, providing a robust validation metric. Our findings indicate that it is possible to effectively reduce the computational cost of the topic modeling process compared to the direct application of LLMs and BERTopic, while simultaneously enhancing inference time and overall efficiency, thereby surpassing the current state-of-the-art capabilities of BERTopic.