<p>Accurate estimation of crop yield over a large extent is crucial for food security, particularly due to the ongoing impact of climate change. This study focused on the integration of machine learning (ML) techniques with multi-source remote sensing data for the estimation of wheat and paddy yield in the Sangrur district of Punjab (north-western India). The decadal yield data (2011–2020) for wheat and paddy, collected from crop cutting experiments, Landsat derived normalized difference vegetation index (NDVI) and climate data (temperature, precipitation, and growing degree days) were used for building the yield estimation models based on three ML algorithms namely random forest regression (RFR), support vector regression (SVR), and gradient boosting regression (GBR). To enhance model performance, a comprehensive hyperparameter tuning process was conducted using randomized search cross-validation to identify the optimal configurations for each algorithm. The study identified key variables influencing wheat and paddy yield using the Permutation Feature Importance (PFI) method with RFR. The most important variables for wheat yield estimation were February minimum temperature and NDVI during March and December. For paddy, September precipitation and NDVI were key predictors. Among the three ML models, the RFR model, when combined with variables selected by PFI analysis, achieved the highest performance (R<sup>2</sup> = 0.807, RMSE = 0.181 t/ha, and MAE = 0.130 t/ha for wheat yield estimation, and R<sup>2</sup> = 0.691, RMSE = 0.552 t/ha, and MAE = 0.429 t/ha for paddy yield estimation). The results provide a basis for well-informed decision making in crop monitoring.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Estimation of wheat and paddy yield in the parts of Indian Punjab by integrating climate and Landsat satellite data using machine learning algorithms

  • Shikha Sharda,
  • Raj Setia,
  • Sumit Kumar,
  • Mohit Arora,
  • Brijendra Pateriya

摘要

Accurate estimation of crop yield over a large extent is crucial for food security, particularly due to the ongoing impact of climate change. This study focused on the integration of machine learning (ML) techniques with multi-source remote sensing data for the estimation of wheat and paddy yield in the Sangrur district of Punjab (north-western India). The decadal yield data (2011–2020) for wheat and paddy, collected from crop cutting experiments, Landsat derived normalized difference vegetation index (NDVI) and climate data (temperature, precipitation, and growing degree days) were used for building the yield estimation models based on three ML algorithms namely random forest regression (RFR), support vector regression (SVR), and gradient boosting regression (GBR). To enhance model performance, a comprehensive hyperparameter tuning process was conducted using randomized search cross-validation to identify the optimal configurations for each algorithm. The study identified key variables influencing wheat and paddy yield using the Permutation Feature Importance (PFI) method with RFR. The most important variables for wheat yield estimation were February minimum temperature and NDVI during March and December. For paddy, September precipitation and NDVI were key predictors. Among the three ML models, the RFR model, when combined with variables selected by PFI analysis, achieved the highest performance (R2 = 0.807, RMSE = 0.181 t/ha, and MAE = 0.130 t/ha for wheat yield estimation, and R2 = 0.691, RMSE = 0.552 t/ha, and MAE = 0.429 t/ha for paddy yield estimation). The results provide a basis for well-informed decision making in crop monitoring.