<p>We introduce a novel dataset for visual recognition systems in retail automation, focusing specifically on fruits and vegetables. The dataset comprises 34 species and 65 varieties, featuring fairly balanced classes and including packed goods in plastic bags. We capture each sample from multiple viewpoints and provide additional annotations, such as object count and total weight. Furthermore, a subset of samples for each class includes segmentation masks. This dataset aims to overcome the limitations of current open-access datasets by providing a more comprehensive and diverse set of training data. A total of 72 annotators collected over 100,000 images of 370,000 objects across multiple shops and cities. Around 9,000 images have manual segmentation masks. To facilitate research in this area, we provide baseline results for zero-shot and supervised classification, instance segmentation, and object counting tasks. We also investigate the impact of packaging and background type on model performance. Ultimately, this dataset is designed to support the development of multitask models for visual recognition in offline retail settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Packed Fruits and Vegetables Visual Classification and Segmentation Benchmark

  • Svetlana Illarionova,
  • Sergey Nesteruk,
  • Tatiana Elina,
  • Sergey Bezzateev,
  • Evgeny Burnaev

摘要

We introduce a novel dataset for visual recognition systems in retail automation, focusing specifically on fruits and vegetables. The dataset comprises 34 species and 65 varieties, featuring fairly balanced classes and including packed goods in plastic bags. We capture each sample from multiple viewpoints and provide additional annotations, such as object count and total weight. Furthermore, a subset of samples for each class includes segmentation masks. This dataset aims to overcome the limitations of current open-access datasets by providing a more comprehensive and diverse set of training data. A total of 72 annotators collected over 100,000 images of 370,000 objects across multiple shops and cities. Around 9,000 images have manual segmentation masks. To facilitate research in this area, we provide baseline results for zero-shot and supervised classification, instance segmentation, and object counting tasks. We also investigate the impact of packaging and background type on model performance. Ultimately, this dataset is designed to support the development of multitask models for visual recognition in offline retail settings.