<p>In the realm of traditional Chinese medicine (TCM) diagnosis, tongue diagnosis experiences a transformative shift toward digitalization and intelligence. Traditional tongue diagnosis has long been criticized for its subjectivity and lack of quantitative standards. Moreover, existing deep learning methods for multi-label tongue image classification have limitations in capturing detailed features and semantic information. To address these challenges, this study introduces a Dual-Branch Label Semantic Embedding Network (DBLSEN) model designed to enhance classification accuracy and feature extraction precision. The proposed model comprises two key components: (1) CNN Attention Branch: which branch utilizes a convolutional neural network to extract features from tongue images. It incorporates a multi-head class-specific residual attention module to strengthen feature representation, ensuring accurate captures of spatial features across different categories; (2) Label Semantic Fusion Branch: which employs a visual patch and label smantic embedding module to obtain patch-level features and label semantic features from tongue images. Multiple visual label semantic fusion modules align these semantic features reducing semantic discrepancies and iteratively generating complex visual feature representations for each label. The features extracted by both branches are ultimately fused in the fusion classification module, predicting multi-label tongue image classification results. Experimental evaluations on the MLTID and Tooth-Marked datasets demonstrate that the proposed method achieves mean average precision scores of 91.84% and 94.28%, respectively—outperforming existing methods. Furthermore, on the ChestX-ray14 dataset, the model achieves an average area under the curve of 83.36%, surpassing the best comparative models by 1.05%. These results underscore the model’s robust generalization capabilities and its potential to significantly advance the field of TCM diagnosis. The source code of our model is publicly available at <a href="https://github.com/emptychallenge/DBLSEN">https://github.com/emptychallenge/DBLSEN</a> to ensure reproducibility and facilitate further research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-label classification of tongue images using label semantic embedding and dual-branch network

  • Xiang Lu,
  • Yue Feng,
  • Xudong Jia,
  • Tao Chen

摘要

In the realm of traditional Chinese medicine (TCM) diagnosis, tongue diagnosis experiences a transformative shift toward digitalization and intelligence. Traditional tongue diagnosis has long been criticized for its subjectivity and lack of quantitative standards. Moreover, existing deep learning methods for multi-label tongue image classification have limitations in capturing detailed features and semantic information. To address these challenges, this study introduces a Dual-Branch Label Semantic Embedding Network (DBLSEN) model designed to enhance classification accuracy and feature extraction precision. The proposed model comprises two key components: (1) CNN Attention Branch: which branch utilizes a convolutional neural network to extract features from tongue images. It incorporates a multi-head class-specific residual attention module to strengthen feature representation, ensuring accurate captures of spatial features across different categories; (2) Label Semantic Fusion Branch: which employs a visual patch and label smantic embedding module to obtain patch-level features and label semantic features from tongue images. Multiple visual label semantic fusion modules align these semantic features reducing semantic discrepancies and iteratively generating complex visual feature representations for each label. The features extracted by both branches are ultimately fused in the fusion classification module, predicting multi-label tongue image classification results. Experimental evaluations on the MLTID and Tooth-Marked datasets demonstrate that the proposed method achieves mean average precision scores of 91.84% and 94.28%, respectively—outperforming existing methods. Furthermore, on the ChestX-ray14 dataset, the model achieves an average area under the curve of 83.36%, surpassing the best comparative models by 1.05%. These results underscore the model’s robust generalization capabilities and its potential to significantly advance the field of TCM diagnosis. The source code of our model is publicly available at https://github.com/emptychallenge/DBLSEN to ensure reproducibility and facilitate further research.