<p>Multimodal recommender systems utilize users’ interaction history and associated multimodal information to recommend items effectively. Although there have been significant advancements in current research, some approaches do not fully capture the complexity of user behavior when interacting with multimodal information. This complexity arises from two factors: the comprehensive influence of a large amount of image and text information related to the items, and the preference differences users exhibit for various item factors across different modal scenarios. Therefore, we propose a novel method named MRFFD (multimodal recommender based on feature fusion and decoupling). This method employs two separate attention networks to extract and integrate the key visual and textual features of items. We then decouple the multimodal features to identify different factors of the items. Finally, in both local scenarios (single-modal) and global scenarios (multimodal), we compute user preference scores for each factor to achieve precise recommendations. Experimental results on two open datasets have demonstrated the effectiveness and superiority of our proposed model. Specifically, MRFFD demonstrated substantial improvements in Precision@20, Recall@20, and NDCG@20 metrics, achieving increases of 4.29%, 2.53%, and 3.93% on the Toy dataset, and 5.78%, 13.24%, and 1.55% on the Yelp dataset, respectively, compared to the baseline models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MRFFD: multimodal recommender based on feature fusion and decoupling

  • Yuchao Ping,
  • Shuqin Wang,
  • Ziyi Yang,
  • Yongquan Dong,
  • Rui Jia

摘要

Multimodal recommender systems utilize users’ interaction history and associated multimodal information to recommend items effectively. Although there have been significant advancements in current research, some approaches do not fully capture the complexity of user behavior when interacting with multimodal information. This complexity arises from two factors: the comprehensive influence of a large amount of image and text information related to the items, and the preference differences users exhibit for various item factors across different modal scenarios. Therefore, we propose a novel method named MRFFD (multimodal recommender based on feature fusion and decoupling). This method employs two separate attention networks to extract and integrate the key visual and textual features of items. We then decouple the multimodal features to identify different factors of the items. Finally, in both local scenarios (single-modal) and global scenarios (multimodal), we compute user preference scores for each factor to achieve precise recommendations. Experimental results on two open datasets have demonstrated the effectiveness and superiority of our proposed model. Specifically, MRFFD demonstrated substantial improvements in Precision@20, Recall@20, and NDCG@20 metrics, achieving increases of 4.29%, 2.53%, and 3.93% on the Toy dataset, and 5.78%, 13.24%, and 1.55% on the Yelp dataset, respectively, compared to the baseline models.