HCVINet: A Multimodal Deep Learning approach for Medicinal Plant Classification Using Visual and Semantic Features
摘要
Proper recognition and classification of medicinal plants is fundamental to pharmacology, botany, and traditional medical applications. However, distinguishing closely related species remains difficult, and traditional methods have not been able to fully address these challenges. In this study, we propose a pioneering deep learning architecture, the Hierarchical Contextual Vision Integration Network (HCVINet), which integrates multi-level image feature extraction with contextual-semantic information using natural language processing (NLP) techniques. HCVINet employs a Hierarchical Feature Extraction Network to capture a broad spectrum of visual features—from elementary textures to complex patterns—and fuses these with semantic data through a Contextual-Correlation Integrated Network (CCINet), enabling the model to utilize both visual and textual information for improved classification decisions. Experiments conducted on benchmark medicinal plant datasets demonstrate that HCVINet achieves a classification accuracy of up to 98.3%, outperforming existing state-of-the-art CNN-based models by an average margin of 4.5%. The model also yields higher retrieval rates and improved harmonic mean scores, validating the effectiveness of combining visual and semantic cues. While HCVINet proves to be a robust tool for medicinal plant classification, its performance may be influenced by the quality of the textual corpus, and occasional mismatches between text and image features can affect classification outcomes. Overall, HCVINet offers a significant advancement in automated plant identification and provides promising implications for research and practical applications.