<p>The Dopamine Receptor D2 (DRD2) is a key target in the treatment of neurological and psychiatric disorders such as schizophrenia, Parkinson’s disease, and addiction. Accurate classification of compounds based on their DRD2 inhibitory activity is essential for advancing neuropharmacological drug discovery. This study presents a Bayesian-optimized XGBoost (BO-XGBoost) framework to enhance the prediction of DRD2 inhibitor activity. A curated dataset of 1056 compounds from ChEMBL was processed using two-dimensional molecular descriptors generated with the Mordred toolkit. After a two-stage feature selection process, 309 informative descriptors were retained for model development. The BO-XGBoost model was trained and validated using tenfold cross-validation and benchmarked against baseline XGBoost, Random Forest, Support Vector Machine, and k-Nearest Neighbors models. BO-XGBoost achieved an accuracy of 89.15%, precision of 90.38%, sensitivity of 87.85%, specificity of 90.48%, and an <i>F</i>1 score of 89.10%, outperforming all comparator models. The model’s reliability was further confirmed through applicability domain analysis and Y-scrambling validation. Feature importance analysis identified key molecular descriptors contributing to DRD2 inhibition. These findings demonstrate the potential of Bayesian-optimized ensemble learning to improve predictive performance in targeted drug discovery applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An interpretable Bayesian-optimized XGBoost framework for neuropsychiatric drug candidate classification

  • Teuku Rizky Noviandy,
  • Ghalieb Mutig Idroes,
  • Irsan Hardi

摘要

The Dopamine Receptor D2 (DRD2) is a key target in the treatment of neurological and psychiatric disorders such as schizophrenia, Parkinson’s disease, and addiction. Accurate classification of compounds based on their DRD2 inhibitory activity is essential for advancing neuropharmacological drug discovery. This study presents a Bayesian-optimized XGBoost (BO-XGBoost) framework to enhance the prediction of DRD2 inhibitor activity. A curated dataset of 1056 compounds from ChEMBL was processed using two-dimensional molecular descriptors generated with the Mordred toolkit. After a two-stage feature selection process, 309 informative descriptors were retained for model development. The BO-XGBoost model was trained and validated using tenfold cross-validation and benchmarked against baseline XGBoost, Random Forest, Support Vector Machine, and k-Nearest Neighbors models. BO-XGBoost achieved an accuracy of 89.15%, precision of 90.38%, sensitivity of 87.85%, specificity of 90.48%, and an F1 score of 89.10%, outperforming all comparator models. The model’s reliability was further confirmed through applicability domain analysis and Y-scrambling validation. Feature importance analysis identified key molecular descriptors contributing to DRD2 inhibition. These findings demonstrate the potential of Bayesian-optimized ensemble learning to improve predictive performance in targeted drug discovery applications.