In this paper, we investigate the effectiveness of machine learning models in predicting the total number of wins an NCAA Division I men’s basketball team will achieve in each season. Using historical data from the 2013 to 2024 seasons, we analyze 24 team-level performance metrics, including points per 100 possessions, effective field goal percentages, and power ratings, to train and evaluate multiple models. We conduct a comparative analysis of five machine learning approaches: decision trees, random forests, neural networks, support vector machines (SVMs), and XGBoost. The results show that regression-based models significantly outperform classification-based ones, with XGBoost delivering the most accurate predictions. It achieved 82.19% test accuracy within a ±1 win margin of error, while the random forest model reached 89.04% accuracy using a more ±4 win margin of error. Feature importance analysis across ensemble models revealed that power rating, offensive efficiency, and defensive field goal percentage were the most influential variables in predicting seasonal success. These findings suggest that machine learning offers a viable framework for realistic season forecasting and could support team managers in setting informed performance goals.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Machine Learning Models for Predicting NCAA Division I Basketball Wins

  • Filip Tjärnhell,
  • Ali Mortada,
  • Wei Lu

摘要

In this paper, we investigate the effectiveness of machine learning models in predicting the total number of wins an NCAA Division I men’s basketball team will achieve in each season. Using historical data from the 2013 to 2024 seasons, we analyze 24 team-level performance metrics, including points per 100 possessions, effective field goal percentages, and power ratings, to train and evaluate multiple models. We conduct a comparative analysis of five machine learning approaches: decision trees, random forests, neural networks, support vector machines (SVMs), and XGBoost. The results show that regression-based models significantly outperform classification-based ones, with XGBoost delivering the most accurate predictions. It achieved 82.19% test accuracy within a ±1 win margin of error, while the random forest model reached 89.04% accuracy using a more ±4 win margin of error. Feature importance analysis across ensemble models revealed that power rating, offensive efficiency, and defensive field goal percentage were the most influential variables in predicting seasonal success. These findings suggest that machine learning offers a viable framework for realistic season forecasting and could support team managers in setting informed performance goals.