Computer Vision (CV), Deep Learning (DL), and Convolutional Neural Network (CNN) are among the recent fields explored for research and development to help specially abled people overcome their problems in their day-to-day life. The main motive of this work primarily concerns increasing the independence among visually impaired individuals. This work utilizes the above-mentioned technologies to propose a solution to be used by visually impaired individuals to recognize the day-to-day items that they may wish to purchase in supermarkets. Machine learning and deep-learning models viz SVM, KNN, AlexNet, VGG-16, and VGG-19 are trained on a self-created dataset of 900 images from 30 classes of day-to-day objects. As these models are to be deployed on edge devices that do not possess a very high computational capability, these models are then pruned to reduce their overall size and computational complexity so that they can run smoothly on the hardware of choice. SVM and KNN have an accuracy of 89.97 and 85.88% respectively. For the deep-learning models - AlexNet, VGG-16 and VGG-19 achieved an accuracy of 99.3, 99.3 and 98.24% respectively without pruning. After pruning, the best accuracy achieved was at 70% sparsity with AlexNet, VGG-16 and VGG-19 achieving an overall accuracy of 99.65, 100 and 100%, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Light-Weight AI Based Model for Grocery Classification

  • R. Sreemathy,
  • Mousami Turuk,
  • Soumya Khurana,
  • Jayashree Jagdale

摘要

Computer Vision (CV), Deep Learning (DL), and Convolutional Neural Network (CNN) are among the recent fields explored for research and development to help specially abled people overcome their problems in their day-to-day life. The main motive of this work primarily concerns increasing the independence among visually impaired individuals. This work utilizes the above-mentioned technologies to propose a solution to be used by visually impaired individuals to recognize the day-to-day items that they may wish to purchase in supermarkets. Machine learning and deep-learning models viz SVM, KNN, AlexNet, VGG-16, and VGG-19 are trained on a self-created dataset of 900 images from 30 classes of day-to-day objects. As these models are to be deployed on edge devices that do not possess a very high computational capability, these models are then pruned to reduce their overall size and computational complexity so that they can run smoothly on the hardware of choice. SVM and KNN have an accuracy of 89.97 and 85.88% respectively. For the deep-learning models - AlexNet, VGG-16 and VGG-19 achieved an accuracy of 99.3, 99.3 and 98.24% respectively without pruning. After pruning, the best accuracy achieved was at 70% sparsity with AlexNet, VGG-16 and VGG-19 achieving an overall accuracy of 99.65, 100 and 100%, respectively.