<p>Image inpainting is an important task in computer vision. In traditional transformer-based image inpainting methods, the attention mechanism computes similarity scores among all elements and allocates attention weights accordingly, requiring significant memory and computational resources. Additionally, when adjacent pixels exhibit large information discrepancies, conventional activation functions tend to discard negative activation values, thereby reducing the accuracy of inpainting results. To address these issues, we propose a novel transformer model called sparse TopK value transformer (STK-former). Our model constructs a sparse matrix using only the TopK elements with the highest attention scores, reducing computational complexity while preserving key region information for global modeling. Additionally, we introduce a dynamic K-value adjustment strategy that optimizes attention allocation based on changes in training loss, balancing computational efficiency and inpainting quality. Furthermore, we employ the Swish activation function to retain some negative feature values, ensuring smooth differentiability across the entire domain and improving gradient flow. The nonlinear characteristics also enhance feature learning capability, enabling high-quality inpainting even when adjacent pixels have large information differences. Experimental results on three public datasets show that compared to existing inpainting algorithms, our method reduces computational costs by 30% while improving inpainting accuracy by 5%, demonstrating its superior performance and lightweight nature. Our model effectively preserves global image information, maintains low computational requirements, and achieves high inpainting quality, providing an efficient and reliable solution for image inpainting tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

STK-former: an efficient transformer with adaptive TopK sparse attention for image inpainting

  • Jianchun Jing,
  • Yuhang Liu,
  • Dian Lei,
  • Junyan Li,
  • Yala Tong

摘要

Image inpainting is an important task in computer vision. In traditional transformer-based image inpainting methods, the attention mechanism computes similarity scores among all elements and allocates attention weights accordingly, requiring significant memory and computational resources. Additionally, when adjacent pixels exhibit large information discrepancies, conventional activation functions tend to discard negative activation values, thereby reducing the accuracy of inpainting results. To address these issues, we propose a novel transformer model called sparse TopK value transformer (STK-former). Our model constructs a sparse matrix using only the TopK elements with the highest attention scores, reducing computational complexity while preserving key region information for global modeling. Additionally, we introduce a dynamic K-value adjustment strategy that optimizes attention allocation based on changes in training loss, balancing computational efficiency and inpainting quality. Furthermore, we employ the Swish activation function to retain some negative feature values, ensuring smooth differentiability across the entire domain and improving gradient flow. The nonlinear characteristics also enhance feature learning capability, enabling high-quality inpainting even when adjacent pixels have large information differences. Experimental results on three public datasets show that compared to existing inpainting algorithms, our method reduces computational costs by 30% while improving inpainting accuracy by 5%, demonstrating its superior performance and lightweight nature. Our model effectively preserves global image information, maintains low computational requirements, and achieves high inpainting quality, providing an efficient and reliable solution for image inpainting tasks.