<p>The rapid growth of high-dimensional data has increased the need for effective feature selection methods that improve predictive efficiency without compromising interpretability. This study presents a comprehensive evaluation of the Sequential Squeeze Feature Selection (SSFS) algorithm, a fairly recent development within wrapper-based methods, and contrasts its performance with established sequential methods, specifically Sequential Forward Selection (SFS), Sequential Backward Selection (SBS), Sequential Floating Forward Selection (SFFS), and Sequential Floating Backward Selection (SFBS). The evaluation adopts a diverse collection of twenty-eight multivariate time series datasets, each with varying numbers of features, observations, and origins. SSFS introduces a bidirectional “squeezing” process that alternates between removing and including features, thereby addressing the nesting and backtracking issues found in traditional wrapper approaches. Empirical findings demonstrate that SSFS, on average, achieves <InlineEquation ID="IEq1"><EquationSource Format="TEX">\(13.283\%\)</EquationSource></InlineEquation> and <InlineEquation ID="IEq2"><EquationSource Format="TEX">\(41.714\%\)</EquationSource></InlineEquation> higher predictive accuracy than the baseline (no feature selection) and competing algorithms, respectively. It also reduces RMSE, MAE, and MAPE by <InlineEquation ID="IEq3"><EquationSource Format="TEX">\(39.543\%\)</EquationSource></InlineEquation>, <InlineEquation ID="IEq4"><EquationSource Format="TEX">\(43.442\%\)</EquationSource></InlineEquation>, and <InlineEquation ID="IEq5"><EquationSource Format="TEX">\(56.536\%\)</EquationSource></InlineEquation> on average. The stability and statistical superiority of the algorithm are confirmed through the Wilcoxon signed-rank test (p value <InlineEquation ID="IEq6"><EquationSource Format="TEX">\(\le 0.001\)</EquationSource></InlineEquation>, with rank-biserial correlations ranging from 0.680 to 0.956 across metrics, indicating large effect sizes rather than marginal statistical significance), highlighting its robustness across varying dimensional and temporal structures. In addition to numeric improvements, SSFS improves understanding by measuring the direct impact of each feature on predictive accuracy, which supports the emerging principles of explainable artificial intelligence (XAI). Unlike post-hoc attribution methods such as SHAP and LIME, which rely on feature independence assumptions violated by temporally autocorrelated inputs, SSFS’s accuracy attribution is computed directly from the selection objective function on temporally ordered validation data, making it structurally appropriate for the time series setting. It offers promising implications for complex fields such as epidemiology, environmental modeling, and econometrics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable feature selection in high-dimensional time series using the sequential squeeze algorithm improves predictive accuracy and interpretability

  • Mahadee Al Mobin,
  • Muhammad Tanveer Islam

摘要

The rapid growth of high-dimensional data has increased the need for effective feature selection methods that improve predictive efficiency without compromising interpretability. This study presents a comprehensive evaluation of the Sequential Squeeze Feature Selection (SSFS) algorithm, a fairly recent development within wrapper-based methods, and contrasts its performance with established sequential methods, specifically Sequential Forward Selection (SFS), Sequential Backward Selection (SBS), Sequential Floating Forward Selection (SFFS), and Sequential Floating Backward Selection (SFBS). The evaluation adopts a diverse collection of twenty-eight multivariate time series datasets, each with varying numbers of features, observations, and origins. SSFS introduces a bidirectional “squeezing” process that alternates between removing and including features, thereby addressing the nesting and backtracking issues found in traditional wrapper approaches. Empirical findings demonstrate that SSFS, on average, achieves \(13.283\%\) and \(41.714\%\) higher predictive accuracy than the baseline (no feature selection) and competing algorithms, respectively. It also reduces RMSE, MAE, and MAPE by \(39.543\%\), \(43.442\%\), and \(56.536\%\) on average. The stability and statistical superiority of the algorithm are confirmed through the Wilcoxon signed-rank test (p value \(\le 0.001\), with rank-biserial correlations ranging from 0.680 to 0.956 across metrics, indicating large effect sizes rather than marginal statistical significance), highlighting its robustness across varying dimensional and temporal structures. In addition to numeric improvements, SSFS improves understanding by measuring the direct impact of each feature on predictive accuracy, which supports the emerging principles of explainable artificial intelligence (XAI). Unlike post-hoc attribution methods such as SHAP and LIME, which rely on feature independence assumptions violated by temporally autocorrelated inputs, SSFS’s accuracy attribution is computed directly from the selection objective function on temporally ordered validation data, making it structurally appropriate for the time series setting. It offers promising implications for complex fields such as epidemiology, environmental modeling, and econometrics.