CMIF: A Cross-Modal Information Fusion Approach for Molecular Property Prediction
摘要
Molecular property prediction is vital in bioinformatics, as accurate and efficient predictions can greatly speed up drug discovery and lower R&D costs. In recent years, the development of deep learning has driven the innovation of data-driven molecular representation learning, which has derived sequence-based, graph-based, and geometry-based molecular representations. Many methods have been proposed to fuse molecular representations of different modalities to improve the performance of molecular property prediction. However, most recent studies have focused on one or two modes and combined them in a simple way, they also neglected the alignment of heterogeneous features. To address these challenges, we propose CMIF, a cross-modal information fusion framework for molecular property prediction, which integrates the three modal representations of sequence, graph, and geometry. The proposed CMIF encodes each modal feature separately through a specific neural network and innovatively employs a cross-attention mechanism for deep feature fusion to avoid the limitations of traditional linear operations. To further solve the problem of heterogeneous feature alignment, CMIF introduces a dual self-supervised strategy of contrast learning and cross-modal matching to strengthen the inter-modal semantic consistency. We evaluated the performance of CMIF on seven molecular datasets, and the results show that our method outperforms baseline models in most cases, validating the effectiveness of CMIF in comprehensively capturing molecular features.