Multi-source data feature fusion and machine learning algorithms for storage age identification of Pericarpium Citri Reticulatae (Chenpi)
摘要
Pericarpium Citri Reticulatae (Chenpi) is valued for its age-dependent quality, yet market fraud necessitates objective identification methods. This study established a multi-source data fusion approach combining Fourier transform infrared (FTIR) spectroscopy and gas chromatography-mass spectrometry (GC-MS) metabolomics with machine learning for Chenpi age classification. Sixty samples spanning three age groups (5, 10, and 15 years) were analyzed. For FTIR data, Savitzky-Golay smoothing combined with multiplicative scatter correction (SG + MSC) was identified as the optimal preprocessing method (silhouette score: 0.7417; CH index: 824.89). For GC-MS data, Log2 transformation yielded the best group separation (score: 2.62) and identified 597 differentially abundant compounds via F-test. Three classification models were evaluated using 5-fold stratified cross-validation. The mid-level data fusion combined with support vector machine (SVM) achieved the highest accuracy of 96.7%, outperforming single-source models. These results have demonstrated that multi-source data fusion significantly enhanced Chenpi age classification, providing a robust approach for Chenpi authentication.