Assessing the accuracy of total antioxidant capacity methods using machine learning approaches
摘要
Accurate estimation of antioxidant capacity is complicated by the interactions between the antioxidant molecules, the differences in redox assay chemistries, and inconsistencies in calibration. To assess the validity of four antioxidant assays—2,2′-azino-bis-3-ethylbenzthiazoline-6-sulphonic acid (ABTS), 1,1-diphenyl-2-picrylhydrazyl (DPPH), Ferric Reducing Antioxidant Power (FRAP), and Total Phenolic Content (TPC)—by integrating assay-specific analysis, supervised machine-learning modelling and residual-based classification of antioxidant standard-mixture responses. Replicate-level data from CB1–CB5 phenolic standard mixtures were used to train and compare Linear Regression (LR), Random Forest (RF) and tuned XGBoost (XGB) models. Models were evaluated using grouped cross-validation and an independent grouped test set to avoid leakage between technical replicates from the same experimental condition. The best-performing model was refitted on the full dataset, and residual deviations between observed and model-predicted gallic acid equivalent (GAE) values were used to assign operational interaction classes. Linear Regression showed weak predictive performance under grouped validation, whereas ensemble models improved prediction accuracy. Tuned XGBoost achieved the best overall performance, with the highest grouped cross-validated R2 and independent test-set R2, and the lowest prediction errors among the models tested. Residual-based classification revealed assay- and mixture-dependent response patterns, with most observations classified as additive and smaller subsets showing positive or negative deviations from model-predicted GAE values. Combining assay-specific analysis, grouped model validation and residual-based interaction classification provides a transparent framework for evaluating antioxidant assay behaviour in defined standard mixtures. The results highlight substantial assay- and composition-dependent variability and support the use of machine-learning models to identify mixture responses that deviate from expected patterns. These model-derived interaction classes should be interpreted as analytical indicators for further investigation rather than direct mechanistic proof of biochemical synergy or antagonism.