School dropout poses a significant threat to personal and social development, limiting employment opportunities and perpetuating economic inequality. This study addresses the challenge of identifying student dropouts by focusing on selecting relevant features or variables. We applied five feature selection techniques: ANOVA, mutual information, sequential forward selection, recursive feature elimination, and the least absolute shrinkage and selection operator, highlighting key sociodemographic features. Fifteen machine learning models, including Decision Trees, Support Vector Machines, and K-nearest neighbors, were trained using these techniques and evaluated through cross-validation. The results showed that the mutual information KNN model achieved the highest precision and F1 score, while Sequential Forward Selection provided optimal results for Decision Trees and Support Vector Machines. These results highlight the critical role of feature selection techniques in identifying the variables that significantly affect the effectiveness of predictive models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Dropout Prediction Models Through Feature Selection Techniques

  • Daniel Domínguez-Gómez,
  • Eduardo Sánchez-Jiménez,
  • Yasmín Hernández,
  • Juan de Dios González Torres,
  • Javier Ortiz-Hernandez

摘要

School dropout poses a significant threat to personal and social development, limiting employment opportunities and perpetuating economic inequality. This study addresses the challenge of identifying student dropouts by focusing on selecting relevant features or variables. We applied five feature selection techniques: ANOVA, mutual information, sequential forward selection, recursive feature elimination, and the least absolute shrinkage and selection operator, highlighting key sociodemographic features. Fifteen machine learning models, including Decision Trees, Support Vector Machines, and K-nearest neighbors, were trained using these techniques and evaluated through cross-validation. The results showed that the mutual information KNN model achieved the highest precision and F1 score, while Sequential Forward Selection provided optimal results for Decision Trees and Support Vector Machines. These results highlight the critical role of feature selection techniques in identifying the variables that significantly affect the effectiveness of predictive models.