<p>Generative deep learning provides new approaches for natural image restoration and structural reconstruction. However, virtually restoring Kizil cave murals remains difficult due to their unique artistic style, complex damage, and the need to preserve semantic consistency. This study proposes a multimodal controlled diffusion model that integrates textual and multi-dimensional visual features for high-precision restoration. The model leverages latent space diffusion for high-quality image generation and introduces structural constraints to improve semantic alignment and controllability. A dynamic feature-adaptive GSC (DFA-GSC) module captures local and global features through multi-scale convolution and an adaptive weight generator, enhancing texture perception. For damaged regions, a conditional matching loss helps refine both texture and structure. Experimental results demonstrate that compared to traditional CNNs and single diffusion models, the proposed method achieves superior performance in both evaluation metrics and visual quality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Virtual restoration method of Kizil Grotto murals based on multimodal controlled diffusion models

  • Guoyan Lv,
  • Huiqin Wang,
  • Ke Wang,
  • Huaidong Zhao,
  • Li Zhao

摘要

Generative deep learning provides new approaches for natural image restoration and structural reconstruction. However, virtually restoring Kizil cave murals remains difficult due to their unique artistic style, complex damage, and the need to preserve semantic consistency. This study proposes a multimodal controlled diffusion model that integrates textual and multi-dimensional visual features for high-precision restoration. The model leverages latent space diffusion for high-quality image generation and introduces structural constraints to improve semantic alignment and controllability. A dynamic feature-adaptive GSC (DFA-GSC) module captures local and global features through multi-scale convolution and an adaptive weight generator, enhancing texture perception. For damaged regions, a conditional matching loss helps refine both texture and structure. Experimental results demonstrate that compared to traditional CNNs and single diffusion models, the proposed method achieves superior performance in both evaluation metrics and visual quality.