In this study, we introduce a novel approach for classifying bread edibility using Vision Transformer-based deep learning model, representing the first application of this technology in this domain. Our method segments bread images into patches, which are linearly projected and transformed into sequence of embeddings. To maintain positional context, we append a positional encoding to the embedding sequence. This sequence is processed through multiple multi-head self-attention (MSA) layers and these MSA mechanism adeptly captures complex spatial relationships, enabling the model to understand interactions between various bread image patches. This advanced spatial feature interpretation significantly enhances the model's edibility classification performance. This study was conducted on a novel dataset comprising 2,520 bread images classified into two categories: Edible and Inedible. To improve classification performance, five data augmentation techniques were employed to synthetically expand the training dataset. The proposed model was tested on various datasets, including our own, existing, and merged datasets. We evaluated multiple Vision Transformer (ViT) models, both with and without pre-trained weights, focusing on ViT_Base_16, ViT_Large_16, and ViT_Huge_14. Experimental results demonstrated that the ViT_Huge_14 model, leveraging pretrained weights and augmented data, achieved an average accuracy of 92%, outperforming current state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision Transformers in Evaluating Bread Edibility

  • D. S. Guru,
  • D. Nandini

摘要

In this study, we introduce a novel approach for classifying bread edibility using Vision Transformer-based deep learning model, representing the first application of this technology in this domain. Our method segments bread images into patches, which are linearly projected and transformed into sequence of embeddings. To maintain positional context, we append a positional encoding to the embedding sequence. This sequence is processed through multiple multi-head self-attention (MSA) layers and these MSA mechanism adeptly captures complex spatial relationships, enabling the model to understand interactions between various bread image patches. This advanced spatial feature interpretation significantly enhances the model's edibility classification performance. This study was conducted on a novel dataset comprising 2,520 bread images classified into two categories: Edible and Inedible. To improve classification performance, five data augmentation techniques were employed to synthetically expand the training dataset. The proposed model was tested on various datasets, including our own, existing, and merged datasets. We evaluated multiple Vision Transformer (ViT) models, both with and without pre-trained weights, focusing on ViT_Base_16, ViT_Large_16, and ViT_Huge_14. Experimental results demonstrated that the ViT_Huge_14 model, leveraging pretrained weights and augmented data, achieved an average accuracy of 92%, outperforming current state-of-the-art approaches.