Multimodal Large Language Models (MLLMs) can perform various food-related tasks with high quality. Notably, high-performance MLLMs, such as GPT-4V, can even estimate caloric content from food images. However, these MLLMs often struggle to accurately recognize volume information, which often leads to errors in calorie estimation. To address this issue, we propose a new MLLM framework called CalorieVoL, designed to enhance the recognition of volume information in food items. By integrating this framework into MLLMs like GPT-4V, we achieved higher scores in terms of MAE and correlation coefficients on Nutrition5k compared to simple MLLMs. Our experiments also showed that the volume-aware recognition improved responses in scenarios where accurate volume estimation is critical.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CalorieVoL: Integrating Volumetric Context Into Multimodal Large Language Models for Image-Based Calorie Estimation

  • Hikaru Tanabe,
  • Keiji Yanai

摘要

Multimodal Large Language Models (MLLMs) can perform various food-related tasks with high quality. Notably, high-performance MLLMs, such as GPT-4V, can even estimate caloric content from food images. However, these MLLMs often struggle to accurately recognize volume information, which often leads to errors in calorie estimation. To address this issue, we propose a new MLLM framework called CalorieVoL, designed to enhance the recognition of volume information in food items. By integrating this framework into MLLMs like GPT-4V, we achieved higher scores in terms of MAE and correlation coefficients on Nutrition5k compared to simple MLLMs. Our experiments also showed that the volume-aware recognition improved responses in scenarios where accurate volume estimation is critical.