In this paper, we propose a novel Multimodal Unified Refinement for Multimedia Recommendation (MUR) framework to address three critical challenges in multimodal recommendation: multimodal domain shift, preference-irrelevant multimodal noise, and incomplete multimodal fusion. The MUR framework incorporates three key components to enhance multimodal features for improved recommendation performance. First, a Multimodal Contrast Layer aligns features with recommendation-specific distributions to mitigate the multimodal domain shift between the pre-trained model and the recommendation task. Second, a Modal Purification Layer focuses on removing impurity noise, such as background elements in images or unrelated text in titles, while preserving key features. Finally, a Dual-View Multimodal Fusion Layer ensures a comprehensive understanding of user preferences and item details by pooling diverse insights from multiple modalities. Through extensive experiments on three public datasets, we demonstrate the effectiveness of the MUR framework to tackle the challenges faced by multimodal recommendation in comparison to existing methods, and pave the way for more accurate and robust recommendations in e-commerce and other domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MUR: Multimodal Unified Refinement for Multimedia Recommendation

  • Taoran Fu,
  • Jiaxuan Cao,
  • Ziqian Wang,
  • Min Yang

摘要

In this paper, we propose a novel Multimodal Unified Refinement for Multimedia Recommendation (MUR) framework to address three critical challenges in multimodal recommendation: multimodal domain shift, preference-irrelevant multimodal noise, and incomplete multimodal fusion. The MUR framework incorporates three key components to enhance multimodal features for improved recommendation performance. First, a Multimodal Contrast Layer aligns features with recommendation-specific distributions to mitigate the multimodal domain shift between the pre-trained model and the recommendation task. Second, a Modal Purification Layer focuses on removing impurity noise, such as background elements in images or unrelated text in titles, while preserving key features. Finally, a Dual-View Multimodal Fusion Layer ensures a comprehensive understanding of user preferences and item details by pooling diverse insights from multiple modalities. Through extensive experiments on three public datasets, we demonstrate the effectiveness of the MUR framework to tackle the challenges faced by multimodal recommendation in comparison to existing methods, and pave the way for more accurate and robust recommendations in e-commerce and other domains.