CAMDiff: a cross-attention guided diffusion framework for digital restoration of degraded art paintings
摘要
Historical paintings such as Thangka, murals, and ink paintings suffer from pigment fading, oxidation, and structural damage, challenging digital heritage preservation. Existing methods rarely preserve semantic consistency, color fidelity, and structural detail simultaneously. We propose CAMDiff, a reference-guided coarse-to-fine framework taking a degraded image and a well-preserved reference as inputs, integrating cross-attention guidance, adversarial training, and Mamba-based diffusion super-resolution. A GoldLeafCrossAttention (GLCA) module preserves gold-leaf metallic luster via gated cross-attention fusion, while a ColorGradientCrossAttention (CGCA) module enhances color and structural clarity using color and gradient priors. The super-resolution stage adopts a Mamba-based residual diffusion model capturing long-range dependencies at linear cost, overcoming the quadratic bottleneck of Transformer-based diffusion. A mixed discriminator loss suppresses pigment-induced color banding. We construct ThangkaDB with 8200 paired degraded-restored images and 20,000 high-resolution images. Experiments show CAMDiff outperforms state-of-the-art methods and generalizes to murals and paintings.