<p>Software defect prediction (SDP) often faces challenges related to heterogeneous software metrics, classifier dependency, and severe class imbalance, which may limit the robustness and generalization of feature selection strategies. This study proposes an adaptive feature selection approach to construct structured and discriminative feature subsets that remain effective across diverse datasets and learning models. The proposed method first measures the relationship between each software metric and the defect label using absolute determination power and then applies an adaptive retention rule to iteratively retain features with stronger discriminative contribution. The evaluation was conducted on multiple public defect datasets using several classical machine learning classifiers. Unlike approaches optimized for specific classifiers, the proposed strategy emphasizes cross-classifier robustness and imbalance-aware evaluation through defect recall and Matthews correlation coefficient. Experimental results show that the proposed method achieves a competitive average MCC of 0.257 and a defect recall of 0.446 compared with baseline approaches, although the statistical tests do not indicate significant superiority. Therefore, the proposed method should be interpreted as a comparable and stable alternative for feature selection under imbalanced SDP conditions. Stability and statistical analyses further indicate that the proposed method maintains comparable performance across dataset-classifier combinations. In addition, feature compactness analysis shows that the performance gains are attributable to efficient, interpretable feature subsets, highlighting the importance of robustness-oriented feature selection in SDP.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust adaptive feature selection for imbalanced software defect prediction

  • Aris Puji Widodo,
  • Prajanto Wahyu Adi,
  • Yeva Fadhilah Ashari,
  • Adhe Setya Pramayoga,
  • Norazlina Binti Khamis

摘要

Software defect prediction (SDP) often faces challenges related to heterogeneous software metrics, classifier dependency, and severe class imbalance, which may limit the robustness and generalization of feature selection strategies. This study proposes an adaptive feature selection approach to construct structured and discriminative feature subsets that remain effective across diverse datasets and learning models. The proposed method first measures the relationship between each software metric and the defect label using absolute determination power and then applies an adaptive retention rule to iteratively retain features with stronger discriminative contribution. The evaluation was conducted on multiple public defect datasets using several classical machine learning classifiers. Unlike approaches optimized for specific classifiers, the proposed strategy emphasizes cross-classifier robustness and imbalance-aware evaluation through defect recall and Matthews correlation coefficient. Experimental results show that the proposed method achieves a competitive average MCC of 0.257 and a defect recall of 0.446 compared with baseline approaches, although the statistical tests do not indicate significant superiority. Therefore, the proposed method should be interpreted as a comparable and stable alternative for feature selection under imbalanced SDP conditions. Stability and statistical analyses further indicate that the proposed method maintains comparable performance across dataset-classifier combinations. In addition, feature compactness analysis shows that the performance gains are attributable to efficient, interpretable feature subsets, highlighting the importance of robustness-oriented feature selection in SDP.