Machine Learning Modeling of Apple Production: Global Insight
摘要
Apple production plays a pivotal role in global agriculture, with significant economic implications for major apple-producing nations such as China, Poland, the United States, and India. Accurate forecasting of apple yields is essential for mitigating economic risks associated with supply fluctuations. This study leveraged advanced statistical and machine learning models—ARIMA, NNAR, XGBoost, and Prophet—to analyze historical production data (1961–2022) from the Food and Agriculture Organization (FAO) and to predict future trends. The dataset was partitioned into training (90%) and testing (10%) sets, with preprocessing steps including missing value imputation and normalization. Model performance was evaluated using root mean square error (RMSE), mean absolute error (MAE), and mean absolute scaled error (MASE) metrics. Results revealed that XGBoost outperformed other models across all countries, achieving the lowest error rates (e.g., RMSE: 2495.77 for China, 818.19 for Poland). By contrast, Prophet exhibited the poorest performance, while ARIMA and NNAR showed intermediate accuracy. The study highlights the robustness of XGBoost in capturing nonlinear patterns and its suitability for apple production forecasting. Forecasts for 2023–2028 indicated stable production in China and the United States, with fluctuations in Poland and India, suggesting varying market dynamics. The projected values for 2028 are expected to be 45,984.32 units for China, 3362.387 units for Poland, 4360.043 units for the United States, and 2289.418 units for India. These findings provide actionable insights for stakeholders to optimize resource allocation, pricing strategies, and risk management. The study underscores the potential of integrating real-time climate and Internet of Things (IoT) data to further enhance predictive accuracy, contributing to sustainable agricultural practices.