<p>Student dropout is one of the major concerns for any university. The existing models for student dropout prediction rely heavily on feature selection techniques to enhance the performance of predictive models. The existing models’ predictions are affected by dimensionality and overfitting, failing to adequately generalize across diverse data sets. This paper presents a novel approach that uses clustering as preprocessing to overcome data ambiguity and inconsistencies. It generates a robust classifier while extracting relevant features. A classifier is then built on this pre-processed data. Classifiers built using this approach enhance the predictive accuracy of student dropouts. This classifier is tested over different machine-learning models with cross-validation for generalization. This generated classifier verified a high F1 score. The proposed work highlights the effectiveness of integrating clustering processes with the modeling process and feature reduction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Clustering based pre-processing for feature reduction and robust student dropout classification

  • Nisha Rani,
  • Venkata Suresh Pachigolla,
  • Akshay Kumar

摘要

Student dropout is one of the major concerns for any university. The existing models for student dropout prediction rely heavily on feature selection techniques to enhance the performance of predictive models. The existing models’ predictions are affected by dimensionality and overfitting, failing to adequately generalize across diverse data sets. This paper presents a novel approach that uses clustering as preprocessing to overcome data ambiguity and inconsistencies. It generates a robust classifier while extracting relevant features. A classifier is then built on this pre-processed data. Classifiers built using this approach enhance the predictive accuracy of student dropouts. This classifier is tested over different machine-learning models with cross-validation for generalization. This generated classifier verified a high F1 score. The proposed work highlights the effectiveness of integrating clustering processes with the modeling process and feature reduction.