A learning-dynamics-driven framework for robust small-object detection in complex field environments: YOLOv11l-AGRI-EMA framework
摘要
Accurate fruit detection under natural field conditions remains challenging due to small object size, occlusion, and complex background variability. While most studies focus on architectural modifications, the role of learning dynamics and training stability remains underexplored. This study proposes a learning-dynamics-driven, stability-aware optimisation framework for fig detection based on the YOLOv11l architecture. With only a modest increase in computational cost, the proposed YOLOv11l-AGRI-EMA model integrates Adam-based optimisation with an Efficient Multi-Scale Attention (EMA) mechanism to enhance feature selectivity and convergence behaviour. Experimental results demonstrate strong performance across multiple metrics, including precision (84.66%), recall (82.14%), F1-score (83.38%), and mAP@0.5–0.95 (65.11%). Multi-seed experiments confirm high reproducibility, while statistical analysis (Wilcoxon test, p < 0.05) validates the significance of the improvements. The model also achieves reliable counting performance (R² = 0.971, MAE = 0.19) and robustness under diverse field conditions. Importantly, the findings provide empirical evidence that performance improvements are not solely dependent on architectural design. Instead, optimisation of learning dynamics emerges as a viable strategy that improves performance while introducing only limited computational overhead. These results suggest that optimisation-oriented training strategies, combined with attention mechanisms, can contribute to improved detection performance in agricultural object detection tasks.