Evaluating domain generalization of product image quality assessment models using AI-generated images
摘要
Assessing the generalization ability of product image quality assessment models remains a critical challenge, particularly when operating under visual conditions that differ from those seen during training. This problem relates to domain generalization, where models are expected to maintain reliable predictions under unseen data distributions. This study investigates the robustness of convolutional neural network (CNN)-based quality assessment models when exposed to AI-generated product images that introduce controlled yet realistic visual variations. A synthetic product image dataset was constructed based on a previously defined visual quality taxonomy, introducing controllable variations that are difficult to capture in real-world data. This design enables the use of synthetic imagery as a structured evaluation benchmark rather than solely for training augmentation. Without retraining or fine-tuning, three pre-trained models including a custom CNN, EfficientNetB0, and MobileNetV2, were evaluated under synthetic distribution shifts. The results show that the custom CNN and EfficientNetB0 achieve strong generalization performance, reaching up to 94.16% accuracy. Confidence intervals and paired significance tests confirm the stability of their predictions, while MobileNetV2 exhibits reduced robustness, particularly for visually adjacent quality classes. The proposed framework is further validated through robustness analysis on curated and unfiltered synthetic data, human perceptual evaluation, and experiments with transformer-based architectures. The findings demonstrate that AI-generated images serve as an effective diagnostic tool for robustness evaluation and failure mode analysis. The framework supports reliable deployment of product image quality assessment models and highlights the value of synthetic data as a structured robustness testing resource in multimedia and e-commerce systems.