Plant-based diets have gained significant attention due to their health benefits and environmental sustainability. Infants are unable to communicate their dietary intake. Fecal metabolites provide an alternative approach to accurately assess consumed foods. This study applied machine learning (ML) techniques-including Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting (GB), and Extreme Gradient Boosting (XGB)-to analyze fecal metabolomic profiles associated with plant-based diets. A total of 159 metabolite features were selected based on the coefficient of variation (CV) criterion for analysis using ML models to predict food intake. For each single food, the dataset was divided into six categories, each corresponding to a specific plant-based food: Almond, Avocado, Barley, Broccoli, Oats, and Walnut. In the “multi-food” approach, samples from all six plant-based foods were combined into a single dataset for prediction. Our results showed that the GB model achieved the highest average accuracy when predicting each of the six datasets separately across all plant-based foods (single food). However, when combining all six datasets into one (multi-food) for prediction, the SVM model provided the best predictive performance. Future research should explore the application of deep learning techniques to gain deeper insights into the data and enhance model performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning on Metabolomic Profiles from Fecal to Identify Plant-Based Food Intake

  • Natnicha Charoenwong,
  • Krerkpon Rattanapoom,
  • Umaporn Uawisetwathana,
  • Awanwee Petchkongkaew,
  • Kritanat Chungnoy,
  • Surasit Uypatchawong,
  • Pokpong Songmuang

摘要

Plant-based diets have gained significant attention due to their health benefits and environmental sustainability. Infants are unable to communicate their dietary intake. Fecal metabolites provide an alternative approach to accurately assess consumed foods. This study applied machine learning (ML) techniques-including Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting (GB), and Extreme Gradient Boosting (XGB)-to analyze fecal metabolomic profiles associated with plant-based diets. A total of 159 metabolite features were selected based on the coefficient of variation (CV) criterion for analysis using ML models to predict food intake. For each single food, the dataset was divided into six categories, each corresponding to a specific plant-based food: Almond, Avocado, Barley, Broccoli, Oats, and Walnut. In the “multi-food” approach, samples from all six plant-based foods were combined into a single dataset for prediction. Our results showed that the GB model achieved the highest average accuracy when predicting each of the six datasets separately across all plant-based foods (single food). However, when combining all six datasets into one (multi-food) for prediction, the SVM model provided the best predictive performance. Future research should explore the application of deep learning techniques to gain deeper insights into the data and enhance model performance.