Dynamic Multimodal Fusion via Meta-Learning Towards Multimodal Recommendation
摘要
In this chapter, using micro-videos as an example, we study the issues involved in multimodal fusion mentioned above. We propose a dynamic multimodal fusion method that treats the multimodal fusion of each item as an independent task. More specifically, we develop a novel meta-learning-based multimodal fusion model, named Meta Multimodal Fusion (MetaMMF), to dynamically integrate multimodal information for micro-video recommendation.