GraESM-FuseDTA: adaptive gated multimodal fusion of graph neural networks and protein language models for robust drug-target affinity prediction
摘要
Accurate drug–target affinity prediction is essential for virtual screening, lead compound prioritization, and drug repositioning, especially under cold-start scenarios involving unseen compounds, unseen targets, or unseen drug–target combinations. However, existing methods still face limitations in transferable molecular representation, contextual protein encoding, statistically reliable performance comparison, and biologically interpretable candidate prioritization. In this study, we propose GraESM-FuseDTA, a graph-enhanced protein language model framework for drug–target affinity prediction. In the drug branch, compound SMILES are converted into molecular graphs and encoded by a residual Graph Isomorphism Network with Edge features, enabling the model to capture atom–bond topology and chemically informative substructures. In the protein branch, a frozen ESM-2 protein language model is used to extract contextual sequence representations without fine-tuning the large protein encoder. To model drug–target compatibility more explicitly, GraESM-FuseDTA employs a gated interaction-aware fusion module that preserves drug features, protein features, element-wise interaction features, difference features, and sample-specific gated mixture features. Experiments on the Davis and KIBA datasets show that GraESM-FuseDTA achieves competitive overall performance and consistent advantages in ranking-oriented and variance-explanation metrics across warm start, drug cold start, target cold start, and strict pair cold start settings. Fold-wise statistical significance tests further support the reliability of most top-ranked metric improvements. Interpretability analyses show that the model identifies functionally relevant target sequence regions within the ESM-visible input segment, adaptively adjusts modality contributions across evaluation scenarios, and forms more organized affinity-related latent representations after gated fusion. In addition, a Yamanishi-based case study with SwissDock molecular docking demonstrates that top-ranked predictions can form structurally plausible binding patterns with favorable calculated affinities. These results suggest that GraESM-FuseDTA provides an effective, interpretable, and practically relevant framework for drug–target affinity prediction.