<p>Preserving Vietnam’s traditional musical instruments requires advanced digital solutions capable of accurately recognizing, modeling, and presenting cultural artifacts. This paper introduces ViTIP-Ext, a unified AI-driven framework for multimodal recognition and interactive dissemination of cultural heritage. The pipeline begins with instance segmentation models (i.e., YOLO), trained to achieve 84% mAP50, to detect and localize instruments from a self-curated dataset. The detected regions are then processed through deep learning classifiers (i.e., CNNs, ViT), fine-tuned and enhanced with a novel CMLC algorithm, yielding an F1-score exceeding 97%. To support contextual knowledge retrieval, an ontology-driven <i>(underlying description logic</i> <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42979_2025_4414_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="27" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mathcal{E}\mathcal{L}\)</EquationSource> </InlineEquation>) chatbot, achieving an F1-score above 85%, enables users to query detailed information about each instrument. Finally, AI-powered 3D reconstruction generates high-fidelity digital models, enabling immersive user interaction. Our framework is deployed through a dedicated web platform and published into the Github. By combining segmentation, classification, semantic reasoning, and 3D interaction in a coherent pipeline, ViTIP-Ext offers a scalable technological approach to preserving and promoting Vietnam’s cultural heritage.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViTIP-Ext: AI-Powered 3D Visualization for Vietnamese Instrument Knowledge Preservation

  • Thanh Ma,
  • Hieu Nguyen,
  • Xuan Nguyen,
  • Ky Nguyen,
  • Minh-Thu Tran-Nguyen

摘要

Preserving Vietnam’s traditional musical instruments requires advanced digital solutions capable of accurately recognizing, modeling, and presenting cultural artifacts. This paper introduces ViTIP-Ext, a unified AI-driven framework for multimodal recognition and interactive dissemination of cultural heritage. The pipeline begins with instance segmentation models (i.e., YOLO), trained to achieve 84% mAP50, to detect and localize instruments from a self-curated dataset. The detected regions are then processed through deep learning classifiers (i.e., CNNs, ViT), fine-tuned and enhanced with a novel CMLC algorithm, yielding an F1-score exceeding 97%. To support contextual knowledge retrieval, an ontology-driven (underlying description logic \(\mathcal{E}\mathcal{L}\) ) chatbot, achieving an F1-score above 85%, enables users to query detailed information about each instrument. Finally, AI-powered 3D reconstruction generates high-fidelity digital models, enabling immersive user interaction. Our framework is deployed through a dedicated web platform and published into the Github. By combining segmentation, classification, semantic reasoning, and 3D interaction in a coherent pipeline, ViTIP-Ext offers a scalable technological approach to preserving and promoting Vietnam’s cultural heritage.