Virtual restoration method of Kizil Grotto murals based on multimodal controlled diffusion models
摘要
Generative deep learning provides new approaches for natural image restoration and structural reconstruction. However, virtually restoring Kizil cave murals remains difficult due to their unique artistic style, complex damage, and the need to preserve semantic consistency. This study proposes a multimodal controlled diffusion model that integrates textual and multi-dimensional visual features for high-precision restoration. The model leverages latent space diffusion for high-quality image generation and introduces structural constraints to improve semantic alignment and controllability. A dynamic feature-adaptive GSC (DFA-GSC) module captures local and global features through multi-scale convolution and an adaptive weight generator, enhancing texture perception. For damaged regions, a conditional matching loss helps refine both texture and structure. Experimental results demonstrate that compared to traditional CNNs and single diffusion models, the proposed method achieves superior performance in both evaluation metrics and visual quality.