<p>This study proposes a novel multi-task Vision Transformer (ViT) framework for simultaneous wheat leaf disease classification and pesticide recommendation. The proposed framework incorporates Grad-CAM interpretability to visualize disease-affected regions and enhance model transparency and trustworthiness. RGB wheat leaf images collected from three public Kaggle datasets were used, covering six disease classes and six pesticide categories. Data augmentation techniques expanded the dataset from 10,720 to 11,408 images. The model was evaluated using fivefold cross-validation with accuracy, precision, recall, F1-score, ROC-AUC, and confusion matrices. Experimental results achieved accuracies of 92% ± 0.01 on original images and 94% ± 0.01 on augmented images, demonstrating stable and effective multi-task learning. Comparative benchmarking with ResNet50 and EfficientNet further confirmed the superior performance of the ViT-based framework. Furthermore, Grad-CAM visualizations improved interpretability by highlighting infected leaf regions relevant to model predictions. The proposed framework demonstrates strong potential for precision agriculture applications by providing accurate disease diagnosis, transparent decision support, and pesticide recommendations. Future work will focus on model optimization and deployment on resource-constrained edge devices.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An interpretable vision transformer framework for joint wheat leaf disease classification and pesticide recommendation

  • Fikadu Berie Adugna

摘要

This study proposes a novel multi-task Vision Transformer (ViT) framework for simultaneous wheat leaf disease classification and pesticide recommendation. The proposed framework incorporates Grad-CAM interpretability to visualize disease-affected regions and enhance model transparency and trustworthiness. RGB wheat leaf images collected from three public Kaggle datasets were used, covering six disease classes and six pesticide categories. Data augmentation techniques expanded the dataset from 10,720 to 11,408 images. The model was evaluated using fivefold cross-validation with accuracy, precision, recall, F1-score, ROC-AUC, and confusion matrices. Experimental results achieved accuracies of 92% ± 0.01 on original images and 94% ± 0.01 on augmented images, demonstrating stable and effective multi-task learning. Comparative benchmarking with ResNet50 and EfficientNet further confirmed the superior performance of the ViT-based framework. Furthermore, Grad-CAM visualizations improved interpretability by highlighting infected leaf regions relevant to model predictions. The proposed framework demonstrates strong potential for precision agriculture applications by providing accurate disease diagnosis, transparent decision support, and pesticide recommendations. Future work will focus on model optimization and deployment on resource-constrained edge devices.