Train–test splits matter more for evaluation than for performance
摘要
In machine learning for QSAR/QSPR, the choice of train–test splitting algorithm alters both a model’s realized performance and the accuracy with which that performance is estimated from the held-out test set. Prior work has established that structure-aware splits such as Kennard–Stone and SPXY produce optimistically biased internal estimates, but these characterizations typically examine one method family at a time. Here we quantify how fifteen splitting strategies from four method families differ across fifteen drug-discovery datasets on two criteria (realized external benchmark performance and performance-estimation bias), and further assess whether splitting strategy affects model-selection quality when multiple model classes are optimized simultaneously. A single, consistent pattern holds across all fifteen datasets and all eleven evaluation measures (RMSE, MAE, MedAE,
Scientific contribution
We provide the first head-to-head comparison of fifteen train–test splitting algorithms from four method families on a common drug-discovery regression benchmark, showing that their effect on the credibility of internal performance estimates is roughly an order of magnitude larger than their effect on realized model performance. We further show that this optimism is general across all eleven evaluation measures, give it a geometric explanation in terms of train–test distributional distance, and separate the splitting problem into two tasks, performance estimation and model selection, that have different optimal choices. On this basis we recommend cluster-shuffle splitting as a calibrated default; whereas prior work described the overoptimism of a single method family (Kennard–Stone), our comparison spans four families and all eleven measures and turns this into a concrete recommendation.