Cluster-based topic modelling has been demonstrated to be an effective method for identifying underlying topics within a corpus of text. Such techniques present the opportunity to leverage powerful embedding models to encode textual semantics in embeddings, before identifying dense clusters of these embeddings as representing topics. However, a key necessity for this process is a prior reduction in the dimensionality of high-dimensional embeddings to ensure efficient clustering. Currently, the UMAP algorithm represents the state-of-art algorithm used in topic modelling algorithms such as Top2Vec for this task. In this work, we investigate a novel paradigm in topic modelling, facilitated by the parametric UMAP algorithm. We propose to investigate how the architectural design of neural networks can contribute to parametric dimesionality reduction, to ensure high-quality topic modelling solutions. We achieve this by implementing a modified transformer-encoder architecture, with novel additional residual connections, into a dimensionality reduction pipeline in the benchmark cluster-based topic model Top2Vec, demonstrating the effectiveness of the addition through an in-depth topic analysis from both a metric and quality perspective. The analysis indicates that incorporating the transformer-encoder architecture for parametric dimensionality reduction in Top2Vec results in an enhancement, as measured by widely accepted topic evaluation metrics, which is further enhanced by the introduction of additional residual connections into the network architecture. Moreover, upon human assessment of the identified topics, it is evident that the proposed transformer-encoder pipelines enhance the granularity of the topic modelling solution, when dealing with a small dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Cluster-Based Topic Models Through Parametric Dimensionality Reduction with Transformer-Encoders

  • Ryan Hodgson,
  • Alexandra I. Cristea,
  • John Graham

摘要

Cluster-based topic modelling has been demonstrated to be an effective method for identifying underlying topics within a corpus of text. Such techniques present the opportunity to leverage powerful embedding models to encode textual semantics in embeddings, before identifying dense clusters of these embeddings as representing topics. However, a key necessity for this process is a prior reduction in the dimensionality of high-dimensional embeddings to ensure efficient clustering. Currently, the UMAP algorithm represents the state-of-art algorithm used in topic modelling algorithms such as Top2Vec for this task. In this work, we investigate a novel paradigm in topic modelling, facilitated by the parametric UMAP algorithm. We propose to investigate how the architectural design of neural networks can contribute to parametric dimesionality reduction, to ensure high-quality topic modelling solutions. We achieve this by implementing a modified transformer-encoder architecture, with novel additional residual connections, into a dimensionality reduction pipeline in the benchmark cluster-based topic model Top2Vec, demonstrating the effectiveness of the addition through an in-depth topic analysis from both a metric and quality perspective. The analysis indicates that incorporating the transformer-encoder architecture for parametric dimensionality reduction in Top2Vec results in an enhancement, as measured by widely accepted topic evaluation metrics, which is further enhanced by the introduction of additional residual connections into the network architecture. Moreover, upon human assessment of the identified topics, it is evident that the proposed transformer-encoder pipelines enhance the granularity of the topic modelling solution, when dealing with a small dataset.