Abstract <p>Spoilage in processed meat products, such as poultry and pork sausages, presents significant challenges for food safety, quality control, and waste reduction. This study presents a machine learning-based framework to classify spoilage intensity levels using sensory, physicochemical, and microbiological features. To overcome limitations caused by small datasets, we applied synthetic data augmentation using a tabular variational autoencoder (TVAE) to generate high-fidelity samples that enhance model generalization. Additionally, traditional oversampling techniques such as SMOTE and ADASYN were employed for comparative purposes and to further address class imbalance issues. Seven machine learning classifiers were evaluated logistic regression, support vector machine, <i>K</i>-nearest neighbors, random forest, gradient boosting, voting classifier, and multilayer perceptron. The best classification performance was achieved when models were trained on GAN-based synthetic data and tested on real samples. For poultry sausage spoilage prediction, the gradient boosting classifier reached the highest accuracy of 97%. For pork sausages, random forest achieved the highest accuracy of 95%. These results confirm the effectiveness of data augmentation in improving predictive robustness. To ensure model transparency, we integrated explainable AI techniques SHAP and LIME into the pipeline. These analyses revealed that sampling time, CO<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11947_2025_3971_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="8" /> </InlineMediaObject> <EquationSource Format="TEX">\(_2\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>2</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> concentration, pH, and microbial species such as <i>Lactobacillus curvatus</i> and <i>Leuconostoc carnosum</i> were among the most influential features in spoilage prediction. The combination of synthetic data generation and interpretable machine learning enables a reliable, scalable, and explainable approach to spoilage classification. This methodology has strong potential for enhancing quality control systems in the meat industry while reducing waste and improving safety along the food supply chain.</p> Graphical abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting Spoilage Intensity Level in Sausage Products Using Explainable Machine Learning and GAN-Based Data Augmentation

  • Volkan Ince,
  • Mohamed Bader-El-Den,
  • Ramazan Esmeli,
  • Lalit Maurya,
  • Omer Faruk Sari

摘要

Abstract

Spoilage in processed meat products, such as poultry and pork sausages, presents significant challenges for food safety, quality control, and waste reduction. This study presents a machine learning-based framework to classify spoilage intensity levels using sensory, physicochemical, and microbiological features. To overcome limitations caused by small datasets, we applied synthetic data augmentation using a tabular variational autoencoder (TVAE) to generate high-fidelity samples that enhance model generalization. Additionally, traditional oversampling techniques such as SMOTE and ADASYN were employed for comparative purposes and to further address class imbalance issues. Seven machine learning classifiers were evaluated logistic regression, support vector machine, K-nearest neighbors, random forest, gradient boosting, voting classifier, and multilayer perceptron. The best classification performance was achieved when models were trained on GAN-based synthetic data and tested on real samples. For poultry sausage spoilage prediction, the gradient boosting classifier reached the highest accuracy of 97%. For pork sausages, random forest achieved the highest accuracy of 95%. These results confirm the effectiveness of data augmentation in improving predictive robustness. To ensure model transparency, we integrated explainable AI techniques SHAP and LIME into the pipeline. These analyses revealed that sampling time, CO \(_2\) 2 concentration, pH, and microbial species such as Lactobacillus curvatus and Leuconostoc carnosum were among the most influential features in spoilage prediction. The combination of synthetic data generation and interpretable machine learning enables a reliable, scalable, and explainable approach to spoilage classification. This methodology has strong potential for enhancing quality control systems in the meat industry while reducing waste and improving safety along the food supply chain.

Graphical abstract