UMAIR-FPS: User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style
摘要
The rapid advancement of high-quality image generation models based on AI has generated a deluge of anime illustrations. Recommending illustrations to users has become a challenge. However, existing anime recommendation systems (RS) have focused on text features but still need to integrate image features. In addition, most multi-modal (MM) RS research is constrained by tightly coupled datasets, limiting its applicability to illustrations RS. We propose the User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style (UMAIR-FPS) to tackle these gaps. In the feature extract phase, for image features, we are the first to combine painting style with semantic features to construct a dual-output image encoder for enhancing representation. For text features, we obtain embeddings based on fine-tuning Sentence-Transformers by incorporating domain knowledge that composes a variety of anime text pairs from multilingual mappings, entity relationships, and term explanation perspectives, respectively. In the MM fusion phase, we novelly propose a user-aware multi-modal contribution measurement mechanism to weight MM features dynamically according to user features at the interaction level and employ the DCN-V2 module to model bounded-degree MM crosses effectively. UMAIR-FPS surpasses the SOTA baselines on large real-world datasets.