DietAI24 as a framework for comprehensive nutrition estimation using multimodal large language models
摘要
Accurate dietary assessment is essential for health research. While smartphone-based food image recognition offers a convenient alternative to traditional methods, existing computer vision approaches struggle with real-world food images and analyze only basic macronutrients, limiting their utility for comprehensive nutritional research.
MethodsWe developed DietAI24, a framework for automated nutrition estimation from food images that combines multimodal large language models (MLLMs) with Retrieval-Augmented Generation (RAG) technology to ground the MLLM’s visual recognition in authoritative nutrition databases rather than relying on the model’s internal knowledge. In our work, we used the Food and Nutrient Database for Dietary Studies (FNDDS) as the authoritative nutrition database. Through this approach, DietAI24 enables accurate nutrient estimation without extensive data collection or model training.
ResultsDietAI24 significantly outperforms existing methods when evaluated against commercial platforms and computer vision baselines using the ASA24 and Nutrition5k datasets. Performance is measured through mean absolute error (MAE). DietAI24 achieves a 63% reduction in MAE for food weight estimation and four key nutrients and food components compared to existing methods when tested on real-world mixed dishes (p < 0.05). Notably, DietAI24 estimates 65 distinct nutrients and food components, far exceeding the basic macronutrient profiles of existing solutions.
ConclusionsDietAI24 demonstrates that integrating MLLMs with RAG and standardized nutrition databases can substantially improve dietary assessment accuracy while enabling comprehensive nutrient analysis. This framework offers a scalable solution for nutrition research and clinical applications, potentially transforming large-scale epidemiological studies and personalized dietary interventions through more accurate and less burdensome dietary data collection.