Cross-Modal Entity Alignment Method Based on Contrastive Learning of Text and Images
摘要
The emergence of cross-modal interactive tasks has created a pressing demand for comprehensive utilization of multimodal knowledge. This challenge has spurred the development of multimodal knowledge graphs, which address these requirements by integrating heterogeneous modality-specific information. Nevertheless, multimodal resource fusion remains a critical task in the construction of multimodal knowledge graphs.To address the limitations of existing multimodal fusion approaches that primarily rely on unimodal textual data while neglecting features from other modalities - resulting in inadequate fine-grained cross-modal interaction - we propose a Fine-Grained Cross-Modal Entity Alignment (FGCMEA) model based on image-text reciprocal contrastive learning. The proposed framework employs: (1) a cross-encoder architecture to guide fine-grained feature learning across modalities, (2) reciprocal contrastive learning to enhance unimodal representation refinement, and (3) explicit alignment constraints between visual and textual entities. Specifically, the model first extracts structural image features and textual representations independently, then aligns cross-modal features through dual encoding to achieve refined unimodal representations, and finally optimizes entity matching via reciprocal contrastive loss functions. Comprehensive evaluations on the Data Structure Course Multimodal (DCM) dataset demonstrate FGCMEA’s superior performance, achieving state-of-the-art accuracy of 27.01% and 32.03%, significantly outperforming existing baseline methods.