A Mamba based vision transformer for fine grained image segmentation of mural figures
摘要
Ancient murals, as valuable cultural heritage, preserve records of historical life and beliefs. To document and preserve these legacies, image segmentation algorithms are essential. This paper proposes a novel framework for fine-grained semantic segmentation of mythological figures and ornaments in murals, featuring an M-ViT Encoder and a Pyramid SCconv Decoder. The M-ViT Encoder incorporates a selective state-space model in the transformer units, mitigating correlations in sequential representations by building long-range dependencies, thus capturing local and global semantic information. The Pyramid SCconv Decoder reduces spatial and channel redundancy by applying separative transformations to encoder outputs, enhancing semantic feedback for fine-grained information separation. Experiments on fine-grained semantic segmentation of colored mural figures demonstrate that the proposed method outperforms other segmentation baselines, achieving 68.42% mIoU and 74.59% mAcc, and delivering state-of-the-art performance.