<p>With the growing prevalence of high-dimensional vector temporal data, accurately recognizing and effectively leveraging inter-sequence associations has become imperative in time series analysis. This paper presents clustered temporal decomposition (CTD), a simultaneous learner of cross-sectional clusters and the temporal structure of such data. CTD achieves this by modeling the ground truth as the product of two low-rank matrices, each independently capturing one of the two structural properties. It accommodates a wide range of temporal and clustering models—including state-of-the-art deep learning methods—to describe the data’s latent spaces. CTD distinguishes itself from existing approaches by allowing the two tasks to mutually influence each other during the learning stage. Under mild conditions, the CTD estimator achieves an explicit non-asymptotic convergence rate to the ground truth with high probability. Simulation experiments show that CTD performs especially well under adverse conditions, such as high observation noise and low clusterability of the true cross-sectional embedding. This demonstrates that CTD draws strength from its ability to access a generalized history, whose unsupervised formation is guided by the clusters. Finally, we confirm CTD’s adaptability to real-world environmental problems by deriving interpretable results on the EPA’s air pollutant dataset. Code examples, additional figures, and proofs for the theoretical results can be found in <a href="https://github.com/h4265172/CTD">https://github.com/h4265172/CTD</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Clustered Temporal Decomposition for High-Dimensional Data

  • Halin Shin,
  • Xiao Wang

摘要

With the growing prevalence of high-dimensional vector temporal data, accurately recognizing and effectively leveraging inter-sequence associations has become imperative in time series analysis. This paper presents clustered temporal decomposition (CTD), a simultaneous learner of cross-sectional clusters and the temporal structure of such data. CTD achieves this by modeling the ground truth as the product of two low-rank matrices, each independently capturing one of the two structural properties. It accommodates a wide range of temporal and clustering models—including state-of-the-art deep learning methods—to describe the data’s latent spaces. CTD distinguishes itself from existing approaches by allowing the two tasks to mutually influence each other during the learning stage. Under mild conditions, the CTD estimator achieves an explicit non-asymptotic convergence rate to the ground truth with high probability. Simulation experiments show that CTD performs especially well under adverse conditions, such as high observation noise and low clusterability of the true cross-sectional embedding. This demonstrates that CTD draws strength from its ability to access a generalized history, whose unsupervised formation is guided by the clusters. Finally, we confirm CTD’s adaptability to real-world environmental problems by deriving interpretable results on the EPA’s air pollutant dataset. Code examples, additional figures, and proofs for the theoretical results can be found in https://github.com/h4265172/CTD.