Due to the prevalence of scale variance in nature images, we propose to use image scale as a self-supervised signal for Masked Image Modeling (MIM). Our method involves selecting random patches from the input image and downsampling them to a low resolution format. Our framework utilizes the latest advances in super-resolution (SR) to design the prediction head, which reconstructs the input from low resolution clues and other patches. Image scale signal also allows our SRMAE to capture scale-invariance feature. Experimental results demonstrate that utilizing scale variance as a self-supervised signal enhances the performance and robustness of SRMAE in various resolution-sensitive visual tasks compared with previous MIM methods. Moreover, our SRMAE outperforms the state-of-the-art supervised method DeriveNet by 1.3% in very low-resolution (VLR) recognition tasks and achieves an accuracy of 74.84% in the recognition of low-resolution facial expressions, surpassing the current leading supervised method FMD by 9.48%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SRMAE: Masked Image Modeling for Scale-Invariant Deep Representations

  • Zhiming Wang,
  • Lin Gu,
  • Feng Lu

摘要

Due to the prevalence of scale variance in nature images, we propose to use image scale as a self-supervised signal for Masked Image Modeling (MIM). Our method involves selecting random patches from the input image and downsampling them to a low resolution format. Our framework utilizes the latest advances in super-resolution (SR) to design the prediction head, which reconstructs the input from low resolution clues and other patches. Image scale signal also allows our SRMAE to capture scale-invariance feature. Experimental results demonstrate that utilizing scale variance as a self-supervised signal enhances the performance and robustness of SRMAE in various resolution-sensitive visual tasks compared with previous MIM methods. Moreover, our SRMAE outperforms the state-of-the-art supervised method DeriveNet by 1.3% in very low-resolution (VLR) recognition tasks and achieves an accuracy of 74.84% in the recognition of low-resolution facial expressions, surpassing the current leading supervised method FMD by 9.48%.