<p>The software defect prediction (SDP) research aims to predict defects early in the software lifecycle. It permits stakeholders to improve software quality, functionality, scalability, reliability, and information security aspects for the targeted software. With the digitalisation of enterprises and processes, its extent has grown since enterprises look for reliable, high-quality software applications. Since most of the Software defect identification is done manually during the development and testing phase, SDP has been an area of research in software engineering. Researchers have been trying SDP modelling using Machine Learning (ML) and Deep Learning (DL) techniques from classified feature metrics and semantic information that can be extracted from the Abstract Syntax Tree (AST) and the software’s source code. Both ML and DL based models suffer from the class imbalance problem, and predictions are biased towards the majority class. This study conducts an empirical comparative analysis of 7 benchmark machine learning (ML) classification models trained on the PROMISE code metrics repository for both balanced and unbalanced classified code metrics and proposes ROS-XGB (Random Oversampling based XGBoost) as the best model for SDP. The experiment indicates that the ROS-XGB-based SDP model outperformed the recent models on evaluation metrics like accuracy, precision, recall, and F1 score. By using the ROS-XGB framework, organisations can optimise human resource allocation, maintain project schedules, and ensure the successful delivery of quality software.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ROS-XGB: a machine learning model for software defect prediction

  • Digvijay Narayan Sharma,
  • Dilip Kumar Yadav

摘要

The software defect prediction (SDP) research aims to predict defects early in the software lifecycle. It permits stakeholders to improve software quality, functionality, scalability, reliability, and information security aspects for the targeted software. With the digitalisation of enterprises and processes, its extent has grown since enterprises look for reliable, high-quality software applications. Since most of the Software defect identification is done manually during the development and testing phase, SDP has been an area of research in software engineering. Researchers have been trying SDP modelling using Machine Learning (ML) and Deep Learning (DL) techniques from classified feature metrics and semantic information that can be extracted from the Abstract Syntax Tree (AST) and the software’s source code. Both ML and DL based models suffer from the class imbalance problem, and predictions are biased towards the majority class. This study conducts an empirical comparative analysis of 7 benchmark machine learning (ML) classification models trained on the PROMISE code metrics repository for both balanced and unbalanced classified code metrics and proposes ROS-XGB (Random Oversampling based XGBoost) as the best model for SDP. The experiment indicates that the ROS-XGB-based SDP model outperformed the recent models on evaluation metrics like accuracy, precision, recall, and F1 score. By using the ROS-XGB framework, organisations can optimise human resource allocation, maintain project schedules, and ensure the successful delivery of quality software.