<p>Deep learning is now central to crop-yield prediction (CYP), but reported gains are too often attributed to architectural novelty rather than to the data regime and evaluation design that actually determine performance. This structured critical review treats CYP as an evidence chain rather than a leaderboard problem. Unlike prior surveys that mainly catalogue architectures or application areas, this review contributes a practical evaluative framework that links architecture choice to data regime, target definition, and validation design, helping practitioners judge model suitability beyond reported accuracy alone. It asks whether each model family matches the structure of the available data, whether explanations are validated rather than merely visualized, and whether published benchmarks support claims of transferability and deployment readiness. Across structured genotype-environment-management tables, satellite time series, UAV phenotyping, and multimodal pipelines, the literature does not support a universally superior architecture. Feedforward networks remain competitive when inputs are already distilled into informative agronomic descriptors; convolutional and recurrent models are strongest when spatial or temporal structure is explicit; and hybrid, attention-based, and multimodal systems are most persuasive when they solve a genuine alignment problem across modalities. The field is more mature in model innovation than in benchmark realism. Reported gains are frequently confounded by heterogeneous targets, leakage-prone splits, weak baselines, limited calibration analysis, sparse external validation, and inconsistent reporting across reviews. Explainability studies face a similar problem: SHAP, LIME, saliency maps, Grad-CAM, and attention mechanisms are useful only when fidelity, stability, and agronomic plausibility are evaluated together. For an AI audience, the central lesson is methodological rather than architectural: trustworthy deployment will depend less on deeper networks alone than on stronger benchmark ecosystems, uncertainty-aware reporting, realistic transfer tests, and clearer multimodal design.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A critical review of deep learning methods for trustworthy crop yield prediction

  • Danison Taremwa,
  • Emmanuel Ahishakiye,
  • Aggrey Obbo,
  • Paul Kategaya Kisozi,
  • Fred Kaggwa

摘要

Deep learning is now central to crop-yield prediction (CYP), but reported gains are too often attributed to architectural novelty rather than to the data regime and evaluation design that actually determine performance. This structured critical review treats CYP as an evidence chain rather than a leaderboard problem. Unlike prior surveys that mainly catalogue architectures or application areas, this review contributes a practical evaluative framework that links architecture choice to data regime, target definition, and validation design, helping practitioners judge model suitability beyond reported accuracy alone. It asks whether each model family matches the structure of the available data, whether explanations are validated rather than merely visualized, and whether published benchmarks support claims of transferability and deployment readiness. Across structured genotype-environment-management tables, satellite time series, UAV phenotyping, and multimodal pipelines, the literature does not support a universally superior architecture. Feedforward networks remain competitive when inputs are already distilled into informative agronomic descriptors; convolutional and recurrent models are strongest when spatial or temporal structure is explicit; and hybrid, attention-based, and multimodal systems are most persuasive when they solve a genuine alignment problem across modalities. The field is more mature in model innovation than in benchmark realism. Reported gains are frequently confounded by heterogeneous targets, leakage-prone splits, weak baselines, limited calibration analysis, sparse external validation, and inconsistent reporting across reviews. Explainability studies face a similar problem: SHAP, LIME, saliency maps, Grad-CAM, and attention mechanisms are useful only when fidelity, stability, and agronomic plausibility are evaluated together. For an AI audience, the central lesson is methodological rather than architectural: trustworthy deployment will depend less on deeper networks alone than on stronger benchmark ecosystems, uncertainty-aware reporting, realistic transfer tests, and clearer multimodal design.