Multi-scale Context-Aware Attention Network with Theme Information for Image Aesthetic Assessment
摘要
In recent years, image aesthetic assessment (IAA) has gradually gained widespread attention in both academia and industry. To address the limitations of existing methods in capturing multi-scale local details and utilizing semantic information, this paper proposes a novel aesthetic evaluation model based on the fusion of visual perception and theme information. First, the model employs a pre-trained Vision Transformer (ViT) to extract global visual features and constructs multi-scale feature representations utilizing intermediate layer outputs. Subsequently, to effectively capture contextual information, a Dynamic Adaptive Contextual Anchor Attention (DACAA) module is proposed, which utilizes multi-branch dynamic convolutional kernels to perform adaptive convolution along horizontal and vertical directions respectively. Next, the Scale Swin Transformer Module (SSTM) is applied to the DACAA-processed features for fine-grained modeling, further enhancing local feature representation capabilities. Finally, to better utilize theme information, a Theme-Guided Attention Fusion module (TGAFM) is proposed to integrate theme features and visual features. Experimental results demonstrate that our method achieves promising performance in aesthetic scoring tasks, outperforming some state-of-the-art methods.