The increasing complexity of software systems has made early defect detection a critical component of software quality assurance. This study explores the development of a novel software defect prediction model utilizing hybrid stacked ensemble learning and advanced feature engineering optimization. The proposed model leverages multiple machine learning classifiers, including Random Forest, Support Vector Machines, Naïve Bayes, and Artificial Neural Networks, integrated through a voting ensemble approach. By optimizing feature selection and preprocessing techniques, the model enhances prediction accuracy and robustness, effectively addressing common challenges such as class imbalance and noisy data. The model was rigorously tested on seven NASA MDP benchmark datasets, demonstrating superior performance to traditional single-model and ensemble methods. The results significantly improve metrics such as PofB20, F1-Score, and ROC-AUC, particularly highlighting the model’s efficacy in the Mozilla dataset. This research underscores the importance of integrating diverse classifiers and robust feature engineering techniques to improve software defect prediction, ultimately contributing to higher software quality and reduced maintenance costs. Future work will focus on refining these methods to address emerging challenges in defect prediction, ensuring the evolution of reliable and efficient software systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Software Defect Prediction with Hybrid Stacked Ensemble Learning and Feature Engineering Optimization

  • Meenakshi Dakuru,
  • Ashish Kande

摘要

The increasing complexity of software systems has made early defect detection a critical component of software quality assurance. This study explores the development of a novel software defect prediction model utilizing hybrid stacked ensemble learning and advanced feature engineering optimization. The proposed model leverages multiple machine learning classifiers, including Random Forest, Support Vector Machines, Naïve Bayes, and Artificial Neural Networks, integrated through a voting ensemble approach. By optimizing feature selection and preprocessing techniques, the model enhances prediction accuracy and robustness, effectively addressing common challenges such as class imbalance and noisy data. The model was rigorously tested on seven NASA MDP benchmark datasets, demonstrating superior performance to traditional single-model and ensemble methods. The results significantly improve metrics such as PofB20, F1-Score, and ROC-AUC, particularly highlighting the model’s efficacy in the Mozilla dataset. This research underscores the importance of integrating diverse classifiers and robust feature engineering techniques to improve software defect prediction, ultimately contributing to higher software quality and reduced maintenance costs. Future work will focus on refining these methods to address emerging challenges in defect prediction, ensuring the evolution of reliable and efficient software systems.