Vision Transformer-Based Transfer Learning Approach for Tomato Maturity Stage Classification
摘要
Tomatoes are considered important vegetables for the agricultural industry due to their widespread cultivation and economic significance. It is essential to determine the optimal harvest timing of tomatoes by accurately determining their maturity stages. Traditionally, tomato maturity assessment was carried out by human visual inspection, which is a labour-intensive practice, but now technologies like computer vision and machine learning can analyse colour, texture and other attributes to determine maturity stages precisely. This study bridges the gap between traditional visual tomato maturity assessment and modern machine learning technology, offering a comparative analysis of proposed Vision Transformer (ViT) based transfer learning model with VGG16 and Convolutional Neural Network (CNN) models for improved accuracy in farming practices. To facilitate this study, a novel tomato dataset consisting of 4500 images was taken. These images were employed for training, validation, and testing of the models. In this study, 600 tomato images were utilised for testing to evaluate the performance of ViT, VGG16 and CNN models. The VGG16 model achieved 95.31%, 93.43%, and 92.83% accuracies in training, validation, and testing, while the CNN model achieved 93.02%, 91.88%, and 88.67% respectively. However, the ViT model attained 96.32%, 94.65%, and 95.5% for training, validation, and testing accuracies. The comparative results demonstrate the efficacy of ViT model and it outperforms the VGG16 and CNN models in terms of higher testing accuracy for the tomato dataset.