Enhanced MobileNet-V2 model with adaptive attention and cross-layer fusion for fine-grained identification of rice diseases and pests for sustainable agriculture
摘要
Amid escalating climate and ecological pressures, green and efficient farming hinges on fast, accurate recognition of rice diseases and pests. Field imagery, however, often contains tiny lesions, motion blur, and cluttered backgrounds, which challenge lightweight models intended for edge deployment. We present an enhanced MobileNet-V2 that integrates a lightweight Super-Resolution (SR) front-end, an Adaptive Selective Attention Module (ASAM), and a Cross-Level Feature Fusion Module (CLFFM). SR restores fine textures in low-quality inputs; ASAM combines channel and spatial cues at multiple depths to highlight discriminative regions; CLFFM fuses shallow details with deep semantics to strengthen small-object perception under clutter. Experiments on a mixed dataset of public and in-field rice images show that our model achieves 94.3% accuracy and a 0.938 F1-score, surpassing mainstream Convolutional Neural Network (CNN) baselines while remaining compact (9.7 MB) and fast (11.2 ms/image). The resulting accuracy–efficiency balance supports real-time, on-device diagnosis, reducing pesticide reliance and environmental burden and providing practical support for sustainable agriculture.