<p>Vulnerability prediction models (VPMs) are statistical machine learning algorithms that are trained to identify vulnerable components in large software systems. Recently, a wide range of software metrics, like the number of dependencies and the size of code between modules, have been evaluated as potential indicators (i.e., <i>features</i>) for building VPMs. Notwithstanding the success achieved by these approaches, none of these models has performed better in vulnerability prediction. This study aims to investigate the use of exemplary data (i.e., <i>Bellwether instances</i>) for vulnerability prediction. Thus, this study explores the impact of <i>Bellwether</i> on VPMs. Specifically, we use <i>n-grams</i> to identify features of vulnerable Java code for improved prediction accuracy. We evaluate our approach on ten Java Android applications extracted from the F-Droid repository. Six machine learning algorithms are used, and the prediction results are evaluated in terms of precision, recall, F-measure, ROC-AUC, and Yuen’s statistical test. The finding indicates that the Bellwether method outperformed the growing portfolio with F-measure values ranging from 18.5 to 94.4% across the studied datasets, respectively. We found that the Decision tree emerged as the best model (AUC value of 0.81) compared with the other classifiers when trained with Bellwether instances. Hence, we recommend the application of Bellwether instances when setting up VPMs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating the use of exemplary data for software vulnerability prediction

  • Patrick Kwaku Kudjo,
  • Solomon Mensah,
  • Ebenezer Owusu,
  • Justice Kwame Appati

摘要

Vulnerability prediction models (VPMs) are statistical machine learning algorithms that are trained to identify vulnerable components in large software systems. Recently, a wide range of software metrics, like the number of dependencies and the size of code between modules, have been evaluated as potential indicators (i.e., features) for building VPMs. Notwithstanding the success achieved by these approaches, none of these models has performed better in vulnerability prediction. This study aims to investigate the use of exemplary data (i.e., Bellwether instances) for vulnerability prediction. Thus, this study explores the impact of Bellwether on VPMs. Specifically, we use n-grams to identify features of vulnerable Java code for improved prediction accuracy. We evaluate our approach on ten Java Android applications extracted from the F-Droid repository. Six machine learning algorithms are used, and the prediction results are evaluated in terms of precision, recall, F-measure, ROC-AUC, and Yuen’s statistical test. The finding indicates that the Bellwether method outperformed the growing portfolio with F-measure values ranging from 18.5 to 94.4% across the studied datasets, respectively. We found that the Decision tree emerged as the best model (AUC value of 0.81) compared with the other classifiers when trained with Bellwether instances. Hence, we recommend the application of Bellwether instances when setting up VPMs.