<p>One challenge of time series data mining is discovering diverse patterns of unknown and varying lengths in time series. To that end, we propose a flexible subsequence clustering framework to determine the appropriate subsequence lengths and segmentation positions adaptively, thereby extracting subsequence patterns of different unknown lengths and representations. In our clustering framework, we propose a subsequence similarity metric to minimize subsequence intra-cluster error while tending to minimize subsequence segmentation numbers. In addition, we also incorporate prior knowledge into our model framework naturally for practical use. Finally, a novel algorithm is adopted to optimize our model. Compared with the SOTA methods, our method achieves an accuracy of 100% in synthetic dataset and 98.37% in real dataset, which largely outperforms other methods. Quantitative and qualitative comparisons in various datasets demonstrate the effectiveness of our method for unknown variable length subsequence clustering.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive variable-length subsequence pattern extraction in time series

  • Ke Zhang,
  • Jiangyong Duan,
  • Tiantian Yang,
  • Lili Guo,
  • Congmin Lv

摘要

One challenge of time series data mining is discovering diverse patterns of unknown and varying lengths in time series. To that end, we propose a flexible subsequence clustering framework to determine the appropriate subsequence lengths and segmentation positions adaptively, thereby extracting subsequence patterns of different unknown lengths and representations. In our clustering framework, we propose a subsequence similarity metric to minimize subsequence intra-cluster error while tending to minimize subsequence segmentation numbers. In addition, we also incorporate prior knowledge into our model framework naturally for practical use. Finally, a novel algorithm is adopted to optimize our model. Compared with the SOTA methods, our method achieves an accuracy of 100% in synthetic dataset and 98.37% in real dataset, which largely outperforms other methods. Quantitative and qualitative comparisons in various datasets demonstrate the effectiveness of our method for unknown variable length subsequence clustering.