<p>Herbaria worldwide have been digitizing their collections to preserve specimens, improve research access, and support biodiversity conservation amid environmental changes. This recent digitization process has revealed that thousands of plants have yet to be appropriately identified or reviewed because of the complex and time-consuming classification and the relatively low number of qualified expert taxonomists. Computer Vision techniques could be promising alternatives for supporting plant identification; however, there is a lack of criteria for designing reliable and representative datasets needed to develop robust classification systems. This often occurs because existing datasets aggregate multiple taxonomic groups with substantial differences among their species, failing to represent the practical realities of plant identification tasks in herbaria. To address this challenge, this work introduces a new database of herbarium specimens exclusively from the Piperaceae Giseke family, accompanied by a series of experiments conducted on this dataset. The Piperaceae, also known as the pepper family, is a large botanical family with many species that are intrinsically complex to identify due to their similarities. We selected 10,503 specimens samples on the <i>species</i>Link repository of 236 Piperaceae species across three genera, collected in Brazil. A comprehensive set of experiments evaluated segmentation, feature extraction, and classification algorithms for the dataset performance as reference values. The best performance combined non-handcrafted features (VGG16 and ViT) and the Multilayer Perceptron classifier. The difficulty in identifying some Piperaceae species is due to their morphological characteristics, which requires that the task be submitted for final review by an expert. We hope the database and our experiments described in this work will benefit the research community.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A database for automatic identification of herbarium specimens in Piperaceae family

  • Alexandre Yuji Kajihara,
  • George Azevedo de Queiroz,
  • Marcelo Galeazzi Caxambú,
  • Luiz Eduardo S. Oliveira,
  • Diego Bertolini,
  • André Luis Schwerz

摘要

Herbaria worldwide have been digitizing their collections to preserve specimens, improve research access, and support biodiversity conservation amid environmental changes. This recent digitization process has revealed that thousands of plants have yet to be appropriately identified or reviewed because of the complex and time-consuming classification and the relatively low number of qualified expert taxonomists. Computer Vision techniques could be promising alternatives for supporting plant identification; however, there is a lack of criteria for designing reliable and representative datasets needed to develop robust classification systems. This often occurs because existing datasets aggregate multiple taxonomic groups with substantial differences among their species, failing to represent the practical realities of plant identification tasks in herbaria. To address this challenge, this work introduces a new database of herbarium specimens exclusively from the Piperaceae Giseke family, accompanied by a series of experiments conducted on this dataset. The Piperaceae, also known as the pepper family, is a large botanical family with many species that are intrinsically complex to identify due to their similarities. We selected 10,503 specimens samples on the speciesLink repository of 236 Piperaceae species across three genera, collected in Brazil. A comprehensive set of experiments evaluated segmentation, feature extraction, and classification algorithms for the dataset performance as reference values. The best performance combined non-handcrafted features (VGG16 and ViT) and the Multilayer Perceptron classifier. The difficulty in identifying some Piperaceae species is due to their morphological characteristics, which requires that the task be submitted for final review by an expert. We hope the database and our experiments described in this work will benefit the research community.