Predictive Analytics in Education: A Comparative Analysis of Machine Learning Models for Predicting Student Performance
摘要
This study evaluates the effectiveness of different predictive models in the educational sector, focusing on their ability to classify and predict student performance in Portuguese and Mathematics. Several datasets of data from a middle school was used for this study. A set of machine learning techniques was applied to predict Portuguese and Math grades respectively. Among the models analysed, the SVM model proved to be superior, achieving accuracy rates of 93% in Portuguese and 79% in Mathematics, surpassing other methods such as XGBoost, AdaBoost, Gradient Boost and Random Forest. The research also explores the impact of student characteristics on academic outcomes, using advanced techniques such as SHAP and permutation feature importance. Critical factors identified include parental education level, tutoring frequency and individual characteristics, which show complex correlations with student grades. These findings highlight the potential of machine learning to provide actionable insights that can improve educational strategies and student support, thereby improving educational outcomes and engagement. Future research will aim to expand the dataset to further validate and refine these findings.