The proposed study involves differentiation of input black tea samples by using NIR (Near Infrared) spectroscopy. These samples are distinguished based on the TAN (Tannic Acid) present in the black tea samples. TAN is considered a major part of feature selection as black tea comprises high concentration of TAN. The experiment is performed under absorbance mode with the wavelength of 900 to 1700 nm. Then the obtained responses are pre-processed using the standard scalar method, where the standardization of data is performed. Moreover, the pre-processed data are fed for feature selection. Further, the feature selected responses are discriminated by employing data reduction techniques based on feature selected attribute such as PCA (Principal Component Analysis), UMAP (Uniform Manifold Approximation and Projection), and KPCA (Kernel Principal Component Analysis). Finally, the effectiveness of the proposed work is evaluated in terms of Silhouette, Calinski-Harabasz and Davies-Bouldin. After clustering, the results estimated that KPCA produced satisfactory outcome with Silhouette coefficient value of 0.64. Thus, this clustering method aids in predicting the accuracy of TAN content in the input samples and is found to be a convenient method for identifying the quality of black tea.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Near Infrared Spectroscopy Based Feature Selection of Tannic Acid for Black Tea Evaluation

  • Angiras Modak,
  • Madhurima Moulick,
  • Runu Banerjee Roy

摘要

The proposed study involves differentiation of input black tea samples by using NIR (Near Infrared) spectroscopy. These samples are distinguished based on the TAN (Tannic Acid) present in the black tea samples. TAN is considered a major part of feature selection as black tea comprises high concentration of TAN. The experiment is performed under absorbance mode with the wavelength of 900 to 1700 nm. Then the obtained responses are pre-processed using the standard scalar method, where the standardization of data is performed. Moreover, the pre-processed data are fed for feature selection. Further, the feature selected responses are discriminated by employing data reduction techniques based on feature selected attribute such as PCA (Principal Component Analysis), UMAP (Uniform Manifold Approximation and Projection), and KPCA (Kernel Principal Component Analysis). Finally, the effectiveness of the proposed work is evaluated in terms of Silhouette, Calinski-Harabasz and Davies-Bouldin. After clustering, the results estimated that KPCA produced satisfactory outcome with Silhouette coefficient value of 0.64. Thus, this clustering method aids in predicting the accuracy of TAN content in the input samples and is found to be a convenient method for identifying the quality of black tea.