<p>Studies on software fault-proneness were typically focused on analysis and prediction, with a few exceptions that used explanatory approach. This paper proposes a novel methodology that, for the first time, utilizes a case-control approach for building explanatory models of software fault-proneness. The files with post-release faults are treated as cases and the other files as controls. The cases and controls are matched by size and prerelease fault-proneness (i.e., Bugfixes) is treated as an exposure. The methodology incorporates software metrics as confounders and, for the first time, considers their interactions. Furthermore, the methodology rigorously handles multicollinearity and uses backward elimination to produce the simplest explanatory models. The odds ratios are quantified using conditional logistic regression which leads to efficient estimates, with tighter confidence intervals. The empirical results, based on three Eclipse releases, showed that while some metrics (i.e., Age, Bugfixes and Developers) consistently affected post-release fault-proneness in two or three releases, the effects of other metrics and interactions were release-specific. Additionally, the first-ever systematic exploration of the generalizability in prior explanatory studies showed that, similarly to our study, they experienced limited generalizability of impactful factors, which is likely due to the complex nature of the software and its development processes. Our results have several practical implications: (1) Simple models with 4–7 significant metrics and interactions can explain post-release fault-proneness; (2) The impactful metrics are simple and easy to collect (e.g., files age, the existence of prerelease faults, and developers count); (3) Due to limited generalizability, release/project-specific explanatory models are necessary.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using case-control study to explain software fault-proneness

  • Yasser Alshehri,
  • Katerina Goseva-Popstojanova

摘要

Studies on software fault-proneness were typically focused on analysis and prediction, with a few exceptions that used explanatory approach. This paper proposes a novel methodology that, for the first time, utilizes a case-control approach for building explanatory models of software fault-proneness. The files with post-release faults are treated as cases and the other files as controls. The cases and controls are matched by size and prerelease fault-proneness (i.e., Bugfixes) is treated as an exposure. The methodology incorporates software metrics as confounders and, for the first time, considers their interactions. Furthermore, the methodology rigorously handles multicollinearity and uses backward elimination to produce the simplest explanatory models. The odds ratios are quantified using conditional logistic regression which leads to efficient estimates, with tighter confidence intervals. The empirical results, based on three Eclipse releases, showed that while some metrics (i.e., Age, Bugfixes and Developers) consistently affected post-release fault-proneness in two or three releases, the effects of other metrics and interactions were release-specific. Additionally, the first-ever systematic exploration of the generalizability in prior explanatory studies showed that, similarly to our study, they experienced limited generalizability of impactful factors, which is likely due to the complex nature of the software and its development processes. Our results have several practical implications: (1) Simple models with 4–7 significant metrics and interactions can explain post-release fault-proneness; (2) The impactful metrics are simple and easy to collect (e.g., files age, the existence of prerelease faults, and developers count); (3) Due to limited generalizability, release/project-specific explanatory models are necessary.