<p>Accurately predicting molecular activity is hindered by activity cliffs (ACs), which are sharp potency changes between highly similar compounds that distort the smoothness assumed by modern quantitative structure–activity relationship (SAR) models and graph neural networks (GNNs). Here we introduce AC awareness (ACA), an inductive bias that reshapes GNN latent spaces to account for these discontinuities. Implemented through an ACA loss combining regression with soft-margin triplet contrastive learning, the method dynamically mines high-value activity-cliff triplets during training and corrects inconsistent neighborhoods in latent space. This process yields progressively fewer cliff violations, more coherent activity gradients, and reduced label incoherence across diverse chemical spaces. Evaluated on 52 datasets spanning low-sample narrow-scaffold series, large mixed-scaffold benchmarks, matched-pair cliff classification, and absorption, distribution, metabolism, excretion and toxicity (ADMET) property prediction, ACA consistently improves predictive accuracy over strong extended-connectivity fingerprint (ECFP) and GNN baselines. The approach generalizes across multiple GNN backbones and remains robust under fixed hyperparameters. These results suggest that ACA provides a principled strategy for enhancing molecular property prediction by aligning latent representations with the nonadditive behavior underlying ACs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Activity-cliff awareness enables robust graph learning for molecular property prediction

  • Chao Cui,
  • Xiaorui Su,
  • Zaixi Zhang,
  • Alejandro Velez-Arce,
  • Jianmin Wang,
  • Xiangcheng Shi,
  • Yanbin Zhang,
  • Jie Wu,
  • Yu Zong Chen,
  • Marinka Zitnik,
  • Wanxiang Shen

摘要

Accurately predicting molecular activity is hindered by activity cliffs (ACs), which are sharp potency changes between highly similar compounds that distort the smoothness assumed by modern quantitative structure–activity relationship (SAR) models and graph neural networks (GNNs). Here we introduce AC awareness (ACA), an inductive bias that reshapes GNN latent spaces to account for these discontinuities. Implemented through an ACA loss combining regression with soft-margin triplet contrastive learning, the method dynamically mines high-value activity-cliff triplets during training and corrects inconsistent neighborhoods in latent space. This process yields progressively fewer cliff violations, more coherent activity gradients, and reduced label incoherence across diverse chemical spaces. Evaluated on 52 datasets spanning low-sample narrow-scaffold series, large mixed-scaffold benchmarks, matched-pair cliff classification, and absorption, distribution, metabolism, excretion and toxicity (ADMET) property prediction, ACA consistently improves predictive accuracy over strong extended-connectivity fingerprint (ECFP) and GNN baselines. The approach generalizes across multiple GNN backbones and remains robust under fixed hyperparameters. These results suggest that ACA provides a principled strategy for enhancing molecular property prediction by aligning latent representations with the nonadditive behavior underlying ACs.