In machine learning (ML) applications, cross-validation (CV) allows greater generalizability of a trained algorithm over out-of-sample or new data. This study explores the accuracy of trained ML algorithms in predicting student performance in a maritime simulator exercise scenario in four different k-fold CVs. Three, five, eight, and ten-fold CVs were trained using a cloud-ML platform. Three top-performing ML algorithms were evaluated considering log loss, accuracy, and area under the curve (AUC). The results indicate higher predictive accuracy with increasing k in CV folds. Considering the trade-off between prediction accuracy and the time required to predict every 1000 observations, using the five-fold CV in predictive learning analytics appears optimal in the explored simulation training scenario. Prediction explanations of five-fold CV are reported.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sensitivity of Predictive Performance Assessment Accuracy in Varying k-fold Cross Validation

  • Fabian Kjeldsberg,
  • Ziaul Haque Munim,
  • Morten Bustgaard,
  • Sahil Bhagat,
  • Emilia Lindroos,
  • Per Haavardtun

摘要

In machine learning (ML) applications, cross-validation (CV) allows greater generalizability of a trained algorithm over out-of-sample or new data. This study explores the accuracy of trained ML algorithms in predicting student performance in a maritime simulator exercise scenario in four different k-fold CVs. Three, five, eight, and ten-fold CVs were trained using a cloud-ML platform. Three top-performing ML algorithms were evaluated considering log loss, accuracy, and area under the curve (AUC). The results indicate higher predictive accuracy with increasing k in CV folds. Considering the trade-off between prediction accuracy and the time required to predict every 1000 observations, using the five-fold CV in predictive learning analytics appears optimal in the explored simulation training scenario. Prediction explanations of five-fold CV are reported.