Emotion-style dual prediction: a multi-task deep learning approach for artistic images
摘要
Artworks convey rich emotional and stylistic information, and jointly predicting artistic style and audience emotion is essential for understanding artistic perception. Traditional methods struggle to capture the abstraction and complexity of art images. To address this, we propose a unified multi-task learning framework integrating Convolutional Neural Networks (CNN) and Vision Transformers (ViT). The CNN module extracts fine-grained local and emotional features, while the ViT models global artistic structures through self-attention. By fusing these complementary representations, the model performs simultaneous style classification and emotion prediction with adaptive loss weighting to enhance collaboration. Experiments on four benchmark datasets—ArtEmis, Affective Image, ADE20K, and Abstract Art—show up to 7.8% higher accuracy and 6.5% improvement in F1-score over state-of-the-art methods, demonstrating the robustness and efficiency of the proposed CNN–ViT multi-task framework.