<p>Chinese Hakka alkaline rice dumplings are an important festival food in southern China during the Dragon Boat Festival. In this study, a fast and non-destructive identification method for Hakka alkaline rice dumplings based on the combination of hyperspectral technology and machine learning is proposed. Using a hyperspectrometer, 2150 bands of the three types of dumplings were acquired, and the noise and baseline drift were eliminated by preprocessing methods such as SG smoothing, SNV scattering correction, and first-order derivatives. In terms of feature engineering, the improved genetic algorithm (GA) achieves 93.1% data compression rate and screens 149 key bands. Among the six machine learning models, SVM, RF, and ANN reach 100% test accuracy on raw data, and the performance of XGBoost and NB models after feature extraction is improved to 100% and 98.33%, respectively. The model efficiency evaluation shows that feature selection reduces the total running time of XGBoost and ANN.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hyperspectral-Driven Machine Learning Feature Optimization for Rapid Identification in Hakka Alkaline Rice Dumplings

  • Zhiwei Wan,
  • Xuewen He,
  • Liping Liu,
  • Lingyue Liu,
  • Ji Zeng

摘要

Chinese Hakka alkaline rice dumplings are an important festival food in southern China during the Dragon Boat Festival. In this study, a fast and non-destructive identification method for Hakka alkaline rice dumplings based on the combination of hyperspectral technology and machine learning is proposed. Using a hyperspectrometer, 2150 bands of the three types of dumplings were acquired, and the noise and baseline drift were eliminated by preprocessing methods such as SG smoothing, SNV scattering correction, and first-order derivatives. In terms of feature engineering, the improved genetic algorithm (GA) achieves 93.1% data compression rate and screens 149 key bands. Among the six machine learning models, SVM, RF, and ANN reach 100% test accuracy on raw data, and the performance of XGBoost and NB models after feature extraction is improved to 100% and 98.33%, respectively. The model efficiency evaluation shows that feature selection reduces the total running time of XGBoost and ANN.