<p>Ancient Chinese architecture (ACA) has rich historical and cultural value, and the preservation of ACA can be enhanced by designing an efficient multi-label ACA image classification model. However, most existing methods mainly focus on convolutional neural networks to learn the local features of ancient architectural images, which cannot retain the positional information of each part of the ancient architecture and fail to leverage the semantic context effectively. This limitation leads to semantic confusion when dealing with similar ancient architectural structures. To solve these problems, we propose a novel coordinate-to-semantic attention network. It explicitly explores positional information and semantic context in an image to optimize the visual feature representation, thereby improving the accuracy of multi-label ACA image classification. Specifically, the coordinate attention mechanism is introduced to process information from different locations and capture long-range dependencies in the image. Subsequently, the semantic attention module realizes the interaction between visual and semantic information by establishing alignments between the two modalities, thereby producing semantically relevant visual representations for each label. Ultimately, the aggregated label features are processed to output the corresponding predictions. Experimental results on the 6-class ACA dataset and other datasets show that the proposed method significantly outperforms existing methods in terms of classification accuracy. Furthermore, the proposed method is of great significance in the protection and research application of ancient architectural heritage.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A coordinate-to-semantic attention network for multi-label ancient Chinese architecture image classification

  • Sulan Zhang,
  • Fei Wang,
  • Huiyuan Zhou,
  • Lihua Hu,
  • Haifeng Yang,
  • Jifu Zhang,
  • Jianghui Cai

摘要

Ancient Chinese architecture (ACA) has rich historical and cultural value, and the preservation of ACA can be enhanced by designing an efficient multi-label ACA image classification model. However, most existing methods mainly focus on convolutional neural networks to learn the local features of ancient architectural images, which cannot retain the positional information of each part of the ancient architecture and fail to leverage the semantic context effectively. This limitation leads to semantic confusion when dealing with similar ancient architectural structures. To solve these problems, we propose a novel coordinate-to-semantic attention network. It explicitly explores positional information and semantic context in an image to optimize the visual feature representation, thereby improving the accuracy of multi-label ACA image classification. Specifically, the coordinate attention mechanism is introduced to process information from different locations and capture long-range dependencies in the image. Subsequently, the semantic attention module realizes the interaction between visual and semantic information by establishing alignments between the two modalities, thereby producing semantically relevant visual representations for each label. Ultimately, the aggregated label features are processed to output the corresponding predictions. Experimental results on the 6-class ACA dataset and other datasets show that the proposed method significantly outperforms existing methods in terms of classification accuracy. Furthermore, the proposed method is of great significance in the protection and research application of ancient architectural heritage.