The emerging field of computer science research that combines data mining and machine learning in sports analytics, has revolutionized performance analysis and predictive modeling. The integration of advanced algorithms poses various challenges and faces a multitude of obstacles. This study aims to develop a real-time system for predicting cricket match results in the Twenty20 format, utilizing ML and statistical techniques. Various ML and statistical methodologies are explored to determine the best prediction results, including the widely employed Random Forest Regressor, Linear Regression, and Decision Tree Regression. The comparative analysis reveals the strengths and weaknesses of each algorithm in the context of cricket score prediction. While Linear Regression provides a baseline understanding of trends, Decision Trees capture complex interactions. Random Forest and SVM offer robustness and non-linearity, but the ensemble method XGBoost emerges as a standout performer due to its high predictive accuracy, feature importance insights, and resistance to overfitting. Presently, T20 cricket match predictions for the first innings predominantly rely on the current run rate, calculated as the runs scored per number of overs bowled. This simplistic approach overlooks critical factors such as wickets fallen, venue considerations, and the outcome of the toss. Moreover, there exists a gap in predicting match outcomes during the second innings. This paper introduces a model that anticipates scores for each inning, employing the XGBoost algorithm in conjunction with the Random Forest Regressor. The proposed model aims to address the limitations of existing predictive models in T20 cricket, enhancing the accuracy and comprehensiveness of match result forecasts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing T20 Cricket Match Prediction Using XGBoost Algorithm and Statistical Techniques: A Comparative Analysis of Various Regression Algorithms

  • Akhil Tyagi,
  • Amandeep Kaur,
  • Aryan Kamboj,
  • Chayandeep Chaulia,
  • Gandharv Mohan,
  • Manpreet Singh

摘要

The emerging field of computer science research that combines data mining and machine learning in sports analytics, has revolutionized performance analysis and predictive modeling. The integration of advanced algorithms poses various challenges and faces a multitude of obstacles. This study aims to develop a real-time system for predicting cricket match results in the Twenty20 format, utilizing ML and statistical techniques. Various ML and statistical methodologies are explored to determine the best prediction results, including the widely employed Random Forest Regressor, Linear Regression, and Decision Tree Regression. The comparative analysis reveals the strengths and weaknesses of each algorithm in the context of cricket score prediction. While Linear Regression provides a baseline understanding of trends, Decision Trees capture complex interactions. Random Forest and SVM offer robustness and non-linearity, but the ensemble method XGBoost emerges as a standout performer due to its high predictive accuracy, feature importance insights, and resistance to overfitting. Presently, T20 cricket match predictions for the first innings predominantly rely on the current run rate, calculated as the runs scored per number of overs bowled. This simplistic approach overlooks critical factors such as wickets fallen, venue considerations, and the outcome of the toss. Moreover, there exists a gap in predicting match outcomes during the second innings. This paper introduces a model that anticipates scores for each inning, employing the XGBoost algorithm in conjunction with the Random Forest Regressor. The proposed model aims to address the limitations of existing predictive models in T20 cricket, enhancing the accuracy and comprehensiveness of match result forecasts.