St-diffnet: Diffusion-based inpainting of dunhuang murals with structural and textural guidance
摘要
Deep learning-based image inpainting methods have gained significant traction in recent years, with diffusion models emerging as a prominent research focus due to their robust generative capabilities and exceptional performance. However, despite advancements, current methods encounter challenges such as underutilized features, inadequate semantic information, and poor global consistency, particularly in complex image scenarios. These challenges hinder their ability to achieve high-precision image inpainting. To address these limitations, this study presents a novel diffusion-based image inpainting network guided by structural and textural features (ST-DiffNet) for the digital restoration of damaged Dunhuang murals. The proposed network integrates a denoising diffusion probabilistic model (DDPM), a dual-branch feature extraction module (DBFE), an attention-based feature fusion module (DAFF), and an image enhancement module (SMSPCNN). Specifically, the DBFE module effectively extracts structural and textural features, providing more accurate support for subsequent feature fusion. The DAFF module introduces an attention mechanism to enhance the integration of multi-source features, thereby improving semantic coherence and contextual understanding. The SMSPCNN module addresses the limitations of diffusion models in restoring visual details by suppressing noise while preserving fine-grained image information. Extensive qualitative and quantitative experiments on three types of Dunhuang mural datasets and the general scene dataset Places2 demonstrate that the proposed method can effectively inpaint damaged murals by restoring both structural and textural features, while ensuring strong semantic integrity and global consistency in the reconstructed content.