<p>Current cultural relic image inpainting methods mainly utilize single encoder-decoder architectures. However, single encoder-decoder methods struggle with introducing prior conditions, especially for historical relics datasets with unique damage patterns. A Multi-column Condition Decoding Transformer for Cultural Relic Image Inpainting (MCDT) is proposed to address above issue. The proposed MCDT model employs multi-column decoders integrated into the Transformer through cross-attention mechanism. Specifically, multi-column decoders consist of three branches: (1) a self-attention branch that decodes the encoded latent features, (2) a ground-truth cross-attention branch that enforces constraints from ground truth data, and (3) an edge cross-attention branch that incorporates edge constraints. The multi-column decoding architecture enables the simultaneous integration of multiple external conditions to constitute a multi-prior constrained image inpainting model. Comparative experiments conducted on cultural relic dataset show that the proposed MCDT method generates higher-quality inpainting results compared to state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cultural Relic Image Inpainting via Multi-column Condition Decoding Transformer

  • Changhong Shi,
  • Weirong Liu,
  • Zhijun Li,
  • Jiajing Yi,
  • Jie Liu

摘要

Current cultural relic image inpainting methods mainly utilize single encoder-decoder architectures. However, single encoder-decoder methods struggle with introducing prior conditions, especially for historical relics datasets with unique damage patterns. A Multi-column Condition Decoding Transformer for Cultural Relic Image Inpainting (MCDT) is proposed to address above issue. The proposed MCDT model employs multi-column decoders integrated into the Transformer through cross-attention mechanism. Specifically, multi-column decoders consist of three branches: (1) a self-attention branch that decodes the encoded latent features, (2) a ground-truth cross-attention branch that enforces constraints from ground truth data, and (3) an edge cross-attention branch that incorporates edge constraints. The multi-column decoding architecture enables the simultaneous integration of multiple external conditions to constitute a multi-prior constrained image inpainting model. Comparative experiments conducted on cultural relic dataset show that the proposed MCDT method generates higher-quality inpainting results compared to state-of-the-art methods.