RoEMF: rotational embedding multimodal fusion for link prediction
摘要
Multimodal link prediction aims to identify missing head and tail entities in the relational triples of multimodal knowledge graphs. However, each modality contains distinct information, and how to effectively fuse multimodal data has become a complex challenge. To address this issue, the rotational embedding multimodal fusion (RoEMF) model was proposed based on rotary position encoding (RoPE). The model employs a multi-head cross-attention mechanism, combined with RoPE, to enhance the representation of positional and contextual information, thereby improving the multimodal data fusion. It focuses on integrating information from different subspaces while capturing cross-modal correlations to mitigate potential data loss, enhances feature fusion, and optimizes the heterogeneity of the representation. Additionally, the cross-modal joint decision loss was proposed to reduce the model’s reliance on single-modal data, aiding in the identification of missing head and tail entities, while enhancing the accuracy and generalization ability of multimodal link prediction. Experimental results on three public MMKG benchmarks demonstrate the outstanding performance of RoEMF compared with other methods in link prediction.