Culturally informed image generation with interpretable latent controls using StyleGAN2
摘要
The growing use of generative artificial intelligence in digital art education has highlighted limitations in existing image synthesis models, particularly their limited support for culturally specific visual patterns and guided learner interaction. While StyleGAN2 delivers high fidelity to the image, it lacks an explicit integration of culturally informed visual features or pedagogically relevant control buttons.
ObjectiveThis paper proposes StyleGAN2 CuHA, a culturally informed generative framework built upon StyleGAN2, to support digital content creation and exploratory learning in public art education.
MethodsThe framework consists of three components: an embedding module for cultural style, a dual style fusion module to control the blending of visual features from the benchmark and cultural style, and a pedagogical latent guidance module to control the blending attributes, including motif density, symmetry, color composition, etc., for users. The mixed dataset training strategy consists of WikiArt data together with selected cultural art datasets, such as TEXMET and BAM. The perceptual metrics, the accuracy of style classification, the motif-level similarity measures, the ablation analysis, and an exploratory user study are used for the evaluation.
ResultsThe results indicate that StyleGAN2 CuHA achieves better FID, LPIPS, style classification accuracy and motif-level similarity than a StyleGAN2 baseline with the same experimental conditions. There are also suggestions from users that improved perceived controllability and visual resemblance would also be desirable.
ConclusionsThe results show that emergence of culturally extended versions of existing generative models is possible to facilitate visually driven exploration and interaction learning in the context of digital art education, but cultural authenticity needs to be validated from a broader geographical community.