Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in performing complex reasoning tasks by leveraging their language-based knowledge, including insights in the food domain. Building on this capability, we hypothesize that MLLMs can enhance calorie estimation from food images by incorporating the language-based reasoning, which is lacking in existing calorie estimation models. However, the effectiveness of these models, particularly when generating text-based outputs for calorie estimation, has not been fully explored. In this work, we present CalorieLLaVA, a model fine-tuned on paired food images and calorie data to exploit the reasoning potential of MLLMs in image-based calorie estimation. By fine-tuning the LLaVA model on the Nutrition5k dataset, we evaluate its performance in calorie estimation. Our experiments demonstrate that CalorieLLaVA surpasses the baseline models, including GPT-4V, GPT-4o, and FoodLMM, achieving superior results on the Nutrition5k dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CalorieLLaVA: Image-Based Calorie Estimation with Multimodal Large Language Models

  • Hikaru Tanabe,
  • Keiji Yanai

摘要

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in performing complex reasoning tasks by leveraging their language-based knowledge, including insights in the food domain. Building on this capability, we hypothesize that MLLMs can enhance calorie estimation from food images by incorporating the language-based reasoning, which is lacking in existing calorie estimation models. However, the effectiveness of these models, particularly when generating text-based outputs for calorie estimation, has not been fully explored. In this work, we present CalorieLLaVA, a model fine-tuned on paired food images and calorie data to exploit the reasoning potential of MLLMs in image-based calorie estimation. By fine-tuning the LLaVA model on the Nutrition5k dataset, we evaluate its performance in calorie estimation. Our experiments demonstrate that CalorieLLaVA surpasses the baseline models, including GPT-4V, GPT-4o, and FoodLMM, achieving superior results on the Nutrition5k dataset.