Multi-view Text Data Stream Clustering via Late Fusion
摘要
Multi-view clustering has recently gained significant attention due to its effective way of leveraging complementary information from multiple representations of the same data compared to single-view methods. However, these methods are primarily designed for static datasets where all views are available in advance, making them unsuitable for handling dynamic, continuously evolving multi-view data streams, where data arrives from multiple sources in real time and the underlying patterns may change over time. To address these limitations, we introduce the MVTStream algorithm, a novel approach designed for clustering multi-view text data streams. Our method operates in three main steps. First, it leverages multiple text representation techniques to generate diverse views of each incoming data stream. Next, these views are clustered in real-time using the concept of micro-clusters. Finally, ensemble methods, such as the Cluster-Based Similarity Partitioning Matrix and the Pairwise Dissimilarity Matrix, are employed to effectively integrate the different views, producing cohesive and accurate clusters. The method demonstrates superior performance over existing techniques in terms of accuracy and adaptability, offering significant advancements in handling dynamic, high-dimensional text streams. Our method not only improves the accuracy and reliability of the clustering process but also offers significant flexibility, making it applicable to a wide range of streaming data scenarios.