Forecasting cotton yield in India using a super ensemble model of machine learning and deep learning techniques
摘要
This paper shows a comprehensive analysis and prediction model for annual cotton yield in India using time series data from 1964 to 2020. We used lag1 to lag5 predictors and a number of machine learning algorithms to predict cotton yield. We have used machine learning algorithms like feedforward neural networks, random forests, and deep learning algorithms like recurrent neural networks and gated recurrent units. Performance metrics like RMSE, MSE, MAPE, MAE and the statistical measures like Standard deviation and mean are evaluated on each model. Initial comparisons are based on root mean square error (RMSE), which identifies lag1 as the most sustainable predictor, yielding the most accurate predictions. We looked at a number of forecasting skill scores, such as accuracy, bias, probability of detection (POD), false alarm ratio (FAR), probability of false detection (POFD), Heidke skill score (HSS), and threat score (TS), to find the most reliable predictor. According to our analysis it shows that GRU demonstrates comparatively better results than other algorithms based on lag1 predictor. Based on the sustainability of the predictor, we have made 10 neural network models and performed ensemble modelling for each configuration like 2, 3, 4, and 5 neurons. After that we created a super ensemble model by combining all the ensemble readings, with the target of improving more accurate predictions. Our analysis demonstrates that the super ensemble model extensively improves the performance of the predictions, which shows the effectiveness of combining multiple models and configurations. This paper gives the valuable insights into the application of machine learning and ensemble techniques in forecasting the agricultural yield, which offers a novelty to forecast the yield of cotton with improved reliability. Additionally, we have applied this methodology to other crops and regions, such as rice crop yield prediction in Lucknow, where the super ensemble model has also demonstrated highly effective results. This extension further validates the robustness and adaptability of our approach across different agricultural domains.