<p>Determining the importance of features and learning feature patterns has proven to be a useful tool for improving the performance of data analysis techniques such as clustering and classification. In this paper, two enhanced versions of the Relief feature weighting algorithm, named C-Relief and V-Relief, are presented to detect sensitive feature interactions and improve the clustering in time series data. In this work, the distribution of samples across clusters is considered for computing the weights and transforming the feature space. The enhanced algorithms compute the diversity of the selected samples based on the clustering labels and propose a wiser weight-updating process. C-Relief refines the convergence of clusters, and V-Relief reduces the number of samples. A novel relabeling method is introduced based on the ensemble clustering concept to adapt clustering for the purpose of prediction in time series data. Historical weather data from Szeged and Berlin datasets, as well as Synthetic Control and OSULeaf data from the UCR classification archive, are evaluated. The enhanced algorithms are compared with other feature weighting algorithms in terms of clustering performance and classification accuracy. C-Relief and V-Relief algorithms and the relabeling method outperform the “PCA, SVM” classifier and are effective in computing the clustering and classification labels of new observations. Evaluation results demonstrate the effectiveness of the proposed algorithms in computing feature weights, producing stable feature patterns, and improving the clustering results.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diversity-Based Feature Weighting to Learn Important Feature Interactions in Time Series Data

  • Ainaz Bahramlou,
  • Zeinab Zali,
  • Massoud Reza Hashemi

摘要

Determining the importance of features and learning feature patterns has proven to be a useful tool for improving the performance of data analysis techniques such as clustering and classification. In this paper, two enhanced versions of the Relief feature weighting algorithm, named C-Relief and V-Relief, are presented to detect sensitive feature interactions and improve the clustering in time series data. In this work, the distribution of samples across clusters is considered for computing the weights and transforming the feature space. The enhanced algorithms compute the diversity of the selected samples based on the clustering labels and propose a wiser weight-updating process. C-Relief refines the convergence of clusters, and V-Relief reduces the number of samples. A novel relabeling method is introduced based on the ensemble clustering concept to adapt clustering for the purpose of prediction in time series data. Historical weather data from Szeged and Berlin datasets, as well as Synthetic Control and OSULeaf data from the UCR classification archive, are evaluated. The enhanced algorithms are compared with other feature weighting algorithms in terms of clustering performance and classification accuracy. C-Relief and V-Relief algorithms and the relabeling method outperform the “PCA, SVM” classifier and are effective in computing the clustering and classification labels of new observations. Evaluation results demonstrate the effectiveness of the proposed algorithms in computing feature weights, producing stable feature patterns, and improving the clustering results.