<p>In recent years, food nutrition estimation has received increasing attention due to its critical role in personalized health management, disease prevention, and scientific dietary guidance. However, traditional physical and chemical methods often rely on manual measurements and expert knowledge, which are both time-consuming and difficult to scale up. With the development of computer vision technology, vision-based food nutrition estimation methods have been proposed, these methods can quickly predict nutritional information in food image without causing damage. However, they still face challenges in fully integrating multi-modal feature and considering the differences between tasks. To solve these issues, we propose a multi-modal fusion network for food nutrition estimation. Specifically, we design a dual-backbone feature extraction architecture. We use Swin Transformer to capture texture and color features from RGB images, while using ConvNeXt to extract spatial structure and volume related features from depth images, ensuring the quality of specific modality features. Then, we use cross attention mechanism to establish fine-grained correlations between RGB and depth features at different scales. To integrate multi-scale information, we introduce the Adaptive Spatial Feature Fusion(ASFF). Furthermore, we fuse ingredient semantic features with visual features through Softmax weighted interaction to enhance feature fusion. In order to solve the differences in estimating different nutrients, we design a task specific regressor with different layers. Extensive experiments are conducted on the Nutrition5K dataset. Our method achieves PMAE values of 12.8% for Calories, 10.1% for Mass, 21.1% for Fat, 18.9% for Carb, and 16.2% for Protein. These results indicate that our outperforms is superior to the existing baseline, validating its effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Task-specialized multi-modal fusion network for food nutrition estimation

  • Shuying Hong,
  • Donglin Zhang,
  • Xiao-Jun Wu

摘要

In recent years, food nutrition estimation has received increasing attention due to its critical role in personalized health management, disease prevention, and scientific dietary guidance. However, traditional physical and chemical methods often rely on manual measurements and expert knowledge, which are both time-consuming and difficult to scale up. With the development of computer vision technology, vision-based food nutrition estimation methods have been proposed, these methods can quickly predict nutritional information in food image without causing damage. However, they still face challenges in fully integrating multi-modal feature and considering the differences between tasks. To solve these issues, we propose a multi-modal fusion network for food nutrition estimation. Specifically, we design a dual-backbone feature extraction architecture. We use Swin Transformer to capture texture and color features from RGB images, while using ConvNeXt to extract spatial structure and volume related features from depth images, ensuring the quality of specific modality features. Then, we use cross attention mechanism to establish fine-grained correlations between RGB and depth features at different scales. To integrate multi-scale information, we introduce the Adaptive Spatial Feature Fusion(ASFF). Furthermore, we fuse ingredient semantic features with visual features through Softmax weighted interaction to enhance feature fusion. In order to solve the differences in estimating different nutrients, we design a task specific regressor with different layers. Extensive experiments are conducted on the Nutrition5K dataset. Our method achieves PMAE values of 12.8% for Calories, 10.1% for Mass, 21.1% for Fat, 18.9% for Carb, and 16.2% for Protein. These results indicate that our outperforms is superior to the existing baseline, validating its effectiveness.