In recent years, image aesthetic assessment (IAA) has gradually gained widespread attention in both academia and industry. To address the limitations of existing methods in capturing multi-scale local details and utilizing semantic information, this paper proposes a novel aesthetic evaluation model based on the fusion of visual perception and theme information. First, the model employs a pre-trained Vision Transformer (ViT) to extract global visual features and constructs multi-scale feature representations utilizing intermediate layer outputs. Subsequently, to effectively capture contextual information, a Dynamic Adaptive Contextual Anchor Attention (DACAA) module is proposed, which utilizes multi-branch dynamic convolutional kernels to perform adaptive convolution along horizontal and vertical directions respectively. Next, the Scale Swin Transformer Module (SSTM) is applied to the DACAA-processed features for fine-grained modeling, further enhancing local feature representation capabilities. Finally, to better utilize theme information, a Theme-Guided Attention Fusion module (TGAFM) is proposed to integrate theme features and visual features. Experimental results demonstrate that our method achieves promising performance in aesthetic scoring tasks, outperforming some state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Context-Aware Attention Network with Theme Information for Image Aesthetic Assessment

  • Wei Li,
  • Chao Long,
  • Zhiqi Liu,
  • Yifa Zheng,
  • Jianfeng Li

摘要

In recent years, image aesthetic assessment (IAA) has gradually gained widespread attention in both academia and industry. To address the limitations of existing methods in capturing multi-scale local details and utilizing semantic information, this paper proposes a novel aesthetic evaluation model based on the fusion of visual perception and theme information. First, the model employs a pre-trained Vision Transformer (ViT) to extract global visual features and constructs multi-scale feature representations utilizing intermediate layer outputs. Subsequently, to effectively capture contextual information, a Dynamic Adaptive Contextual Anchor Attention (DACAA) module is proposed, which utilizes multi-branch dynamic convolutional kernels to perform adaptive convolution along horizontal and vertical directions respectively. Next, the Scale Swin Transformer Module (SSTM) is applied to the DACAA-processed features for fine-grained modeling, further enhancing local feature representation capabilities. Finally, to better utilize theme information, a Theme-Guided Attention Fusion module (TGAFM) is proposed to integrate theme features and visual features. Experimental results demonstrate that our method achieves promising performance in aesthetic scoring tasks, outperforming some state-of-the-art methods.