Background <p>Structure-activity relationship (SAR) modeling of natural products presents a persistent methodological challenge: datasets typically contain fewer than 100 compounds, which restricts the use of data-hungry deep learning models, while conventional QSAR approaches lack mechanistic interpretability and function as black boxes.</p> Methods <p>We present MK-Ensemble, a systematic four-stage optimization framework for interpretable, fragment-based SAR modeling with small-sample natural product datasets. The framework integrates multi-kernel support vector regression, hybrid molecular representations, adversarial domain adaptation, and stacking ensemble learning with strictly nested cross-validation. As a validation case study, we applied the framework to a curated dataset of 91 antioxidant compounds–comprising 24 steroidal saponins from <i>Polygonatum cyrtonema</i> and 67 structurally diverse reference compounds–with 128 activity records across DPPH, ABTS, and FRAP assays. We additionally performed applicability domain characterization via Williams plots and descriptor-space distance analysis, Y-randomization testing (500 permutations), and rigorous statistical model comparison using corrected resampled t-tests and Bayesian correlated t-tests.</p> Results <p>The Stacking Ensemble achieved <InlineEquation ID="IEq1"><EquationSource Format="TEX">\(R^2 = 0.846\)</EquationSource><EquationSource Format="MATHML"><math><mrow><msup><mi>R</mi><mn>2</mn></msup><mo>=</mo><mn>0.846</mn></mrow></math></EquationSource></InlineEquation> (95% CI 0.78–0.91; RMSE = 0.154; <InlineEquation ID="IEq2"><EquationSource Format="TEX">\(Q^2_{\textrm{CV}} = 0.831\)</EquationSource><EquationSource Format="MATHML"><math><mrow><msubsup><mi>Q</mi><mtext>CV</mtext><mn>2</mn></msubsup><mo>=</mo><mn>0.831</mn></mrow></math></EquationSource></InlineEquation>) for DPPH and <InlineEquation ID="IEq3"><EquationSource Format="TEX">\(R^2 = 0.920\)</EquationSource><EquationSource Format="MATHML"><math><mrow><msup><mi>R</mi><mn>2</mn></msup><mo>=</mo><mn>0.920</mn></mrow></math></EquationSource></InlineEquation> (95% CI 0.87–0.97; RMSE = 0.089; <InlineEquation ID="IEq4"><EquationSource Format="TEX">\(Q^2_{\textrm{CV}} = 0.907\)</EquationSource><EquationSource Format="MATHML"><math><mrow><msubsup><mi>Q</mi><mtext>CV</mtext><mn>2</mn></msubsup><mo>=</mo><mn>0.907</mn></mrow></math></EquationSource></InlineEquation>) for ABTS, with corrected resampled t-tests confirming significance over Random Forest baselines (DPPH: <InlineEquation ID="IEq5"><EquationSource Format="TEX">\(p = 0.038\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi>p</mi><mo>=</mo><mn>0.038</mn></mrow></math></EquationSource></InlineEquation>; ABTS: <InlineEquation ID="IEq6"><EquationSource Format="TEX">\(p = 0.003\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi>p</mi><mo>=</mo><mn>0.003</mn></mrow></math></EquationSource></InlineEquation>). Ablation experiments confirmed that each optimization stage–fragment feature integration (+ 20%), domain adaptation (+ 17%), and ensemble stacking (+ 32%)–contributes measurably to final performance. Y-randomization testing (500 permutations) demonstrated that observed performance is extremely unlikely under null label distributions (<InlineEquation ID="IEq7"><EquationSource Format="TEX">\(p &lt; 0.002\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi>p</mi><mo>&lt;</mo><mn>0.002</mn></mrow></math></EquationSource></InlineEquation>). Applicability domain analysis via Williams plots confirmed that the majority of test-set predictions fall within the reliable prediction domain. Integrated Gradients (IG) analysis correctly recovered the well-established SAR pattern that aglycone cores contribute substantially more to predicted antioxidant activity than glycosylated fragments (scores 1.63–1.72 vs. 0.41–0.54, Wilcoxon rank-sum <InlineEquation ID="IEq8"><EquationSource Format="TEX">\(p &lt; 0.001\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi>p</mi><mo>&lt;</mo><mn>0.001</mn></mrow></math></EquationSource></InlineEquation>). Attribution stability was validated across 100 bootstrap replicates (Spearman <InlineEquation ID="IEq9"><EquationSource Format="TEX">\(\rho = 0.87 \pm 0.06\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi>ρ</mi><mo>=</mo><mn>0.87</mn><mo>±</mo><mn>0.06</mn></mrow></math></EquationSource></InlineEquation>), and convergent evidence was obtained from SHAP analysis (Spearman <InlineEquation ID="IEq10"><EquationSource Format="TEX">\(\rho = 0.91\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi>ρ</mi><mo>=</mo><mn>0.91</mn></mrow></math></EquationSource></InlineEquation>) and permutation importance testing.</p> Conclusions <p>MK-Ensemble provides an interpretable SAR modeling framework for small-sample natural product datasets. The framework demonstrates that predictive performance and fragment-level interpretability can be obtained concurrently under data scarcity. Computational network pharmacology and molecular dynamics simulations provide complementary multi-scale support for the fragment-level interpretations, though experimental validation remains necessary for definitive mechanistic conclusions.</p> Scientific contribution <p>Current cheminformatics approaches typically treat kernel-based prediction and fragment-based explanation as separate problems; MK-Ensemble unifies multi-kernel learning, fragment attention, adversarial domain adaptation, and stacking ensemble learning into a single interpretable pipeline optimized for small-sample natural product datasets. The framework shows that strong predictive performance and fragment-level mechanistic interpretability can be obtained concurrently when fewer than 100 compounds are available, a regime where deep learning models often struggle. Applied to steroidal saponins, the model recovers the established structure-activity relationship that aglycone cores dominate antioxidant activity over glycosylated fragments, and this computational finding is complemented by network pharmacology and molecular dynamics simulations, which provide multi-scale computational support for the fragment-level interpretations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MK-ensemble: fragment-based multi-kernel ensemble for interpretable structure-activity relationship modeling of steroidal saponins

  • Guohao Lv,
  • Yingchun Xia,
  • Huichao Liu,
  • Xiaolei Zhu,
  • Shuai Yang,
  • Qingyong Wang,
  • Lichuan Gu

摘要

Background

Structure-activity relationship (SAR) modeling of natural products presents a persistent methodological challenge: datasets typically contain fewer than 100 compounds, which restricts the use of data-hungry deep learning models, while conventional QSAR approaches lack mechanistic interpretability and function as black boxes.

Methods

We present MK-Ensemble, a systematic four-stage optimization framework for interpretable, fragment-based SAR modeling with small-sample natural product datasets. The framework integrates multi-kernel support vector regression, hybrid molecular representations, adversarial domain adaptation, and stacking ensemble learning with strictly nested cross-validation. As a validation case study, we applied the framework to a curated dataset of 91 antioxidant compounds–comprising 24 steroidal saponins from Polygonatum cyrtonema and 67 structurally diverse reference compounds–with 128 activity records across DPPH, ABTS, and FRAP assays. We additionally performed applicability domain characterization via Williams plots and descriptor-space distance analysis, Y-randomization testing (500 permutations), and rigorous statistical model comparison using corrected resampled t-tests and Bayesian correlated t-tests.

Results

The Stacking Ensemble achieved \(R^2 = 0.846\)R2=0.846 (95% CI 0.78–0.91; RMSE = 0.154; \(Q^2_{\textrm{CV}} = 0.831\)QCV2=0.831) for DPPH and \(R^2 = 0.920\)R2=0.920 (95% CI 0.87–0.97; RMSE = 0.089; \(Q^2_{\textrm{CV}} = 0.907\)QCV2=0.907) for ABTS, with corrected resampled t-tests confirming significance over Random Forest baselines (DPPH: \(p = 0.038\)p=0.038; ABTS: \(p = 0.003\)p=0.003). Ablation experiments confirmed that each optimization stage–fragment feature integration (+ 20%), domain adaptation (+ 17%), and ensemble stacking (+ 32%)–contributes measurably to final performance. Y-randomization testing (500 permutations) demonstrated that observed performance is extremely unlikely under null label distributions (\(p < 0.002\)p<0.002). Applicability domain analysis via Williams plots confirmed that the majority of test-set predictions fall within the reliable prediction domain. Integrated Gradients (IG) analysis correctly recovered the well-established SAR pattern that aglycone cores contribute substantially more to predicted antioxidant activity than glycosylated fragments (scores 1.63–1.72 vs. 0.41–0.54, Wilcoxon rank-sum \(p < 0.001\)p<0.001). Attribution stability was validated across 100 bootstrap replicates (Spearman \(\rho = 0.87 \pm 0.06\)ρ=0.87±0.06), and convergent evidence was obtained from SHAP analysis (Spearman \(\rho = 0.91\)ρ=0.91) and permutation importance testing.

Conclusions

MK-Ensemble provides an interpretable SAR modeling framework for small-sample natural product datasets. The framework demonstrates that predictive performance and fragment-level interpretability can be obtained concurrently under data scarcity. Computational network pharmacology and molecular dynamics simulations provide complementary multi-scale support for the fragment-level interpretations, though experimental validation remains necessary for definitive mechanistic conclusions.

Scientific contribution

Current cheminformatics approaches typically treat kernel-based prediction and fragment-based explanation as separate problems; MK-Ensemble unifies multi-kernel learning, fragment attention, adversarial domain adaptation, and stacking ensemble learning into a single interpretable pipeline optimized for small-sample natural product datasets. The framework shows that strong predictive performance and fragment-level mechanistic interpretability can be obtained concurrently when fewer than 100 compounds are available, a regime where deep learning models often struggle. Applied to steroidal saponins, the model recovers the established structure-activity relationship that aglycone cores dominate antioxidant activity over glycosylated fragments, and this computational finding is complemented by network pharmacology and molecular dynamics simulations, which provide multi-scale computational support for the fragment-level interpretations.