Multi-modal Entity Alignment (MMEA) aims to establish correlations between modalities such as images and texts to align equivalent entities across different multi-modal knowledge graphs, thereby enhancing knowledge graph coverage and addressing issues of information loss and low coverage in multi-modal knowledge graphs. Existing MMEA techniques mainly focus on heuristic merging paradigms of single-modal embedding. However, due to modality heterogeneity and the absence of visual imagery, current MMEA approaches encounter challenges such as imbalance and ambiguity in multi-modal data fusion, leading to semantic inconsistencies. To address this issue, this paper proposes a feature-balanced Multi-modal Entity Alignment method (FBMEA) and designs a corresponding fusion framework. Different modalities’ information is independently encoded, and through adaptive feature fusion and multi-head attention mechanisms, the training effects of weak modalities like visual information are dynamically adjusted to enhance the utilization of long-tail entities. Experimental results on three public bilingual datasets and two cross-graph datasets demonstrate that the model’s alignment capability surpasses that of current mainstream models, validating the feasibility and effectiveness of FBMEA.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature Balance Method for Multi-modal Entity Alignment

  • Wei Chen,
  • Xiaofei Li,
  • Sheng Long,
  • Jun Lei,
  • Shuohao Li,
  • Jun Zhang

摘要

Multi-modal Entity Alignment (MMEA) aims to establish correlations between modalities such as images and texts to align equivalent entities across different multi-modal knowledge graphs, thereby enhancing knowledge graph coverage and addressing issues of information loss and low coverage in multi-modal knowledge graphs. Existing MMEA techniques mainly focus on heuristic merging paradigms of single-modal embedding. However, due to modality heterogeneity and the absence of visual imagery, current MMEA approaches encounter challenges such as imbalance and ambiguity in multi-modal data fusion, leading to semantic inconsistencies. To address this issue, this paper proposes a feature-balanced Multi-modal Entity Alignment method (FBMEA) and designs a corresponding fusion framework. Different modalities’ information is independently encoded, and through adaptive feature fusion and multi-head attention mechanisms, the training effects of weak modalities like visual information are dynamically adjusted to enhance the utilization of long-tail entities. Experimental results on three public bilingual datasets and two cross-graph datasets demonstrate that the model’s alignment capability surpasses that of current mainstream models, validating the feasibility and effectiveness of FBMEA.