This study evaluates models on the 30vnfood dataset to optimize Vietnamese food captioning. We propose a pipeline combining CLIP and MBart, achieving high BLEU (0.9313), CIDEr (7.4063), and METEOR (0.9311) scores. Integrating vision and language features proved essential for generating accurate captions. Our results provide actionable insights for food e-commerce and VNLP by addressing the challenges of a low-resource language in image captioning tasks. The evaluation of model performance on the Vietnamese food dataset, 30vnfood, is a critical task in determining optimal models for specific objectives. In this study, we propose a novel pipeline combining CLIP-based image embeddings with a trainable MBart model for Vietnamese caption generation. The pipeline utilizes a frozen CLIP model for feature extraction, a linear projection layer for dimensionality alignment, and MBart for natural language generation. This research systematically evaluates the effectiveness of the proposed architecture in generating captions for Vietnamese food images. Our results provide insights into the model’s performance, highlighting its potential for broader applications in food-related industries and automated systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Domain-Specific Image Captioning: Vietnamese Cuisine on the 30VNFoods Dataset

  • Huynh Vu Nhu Nguyen,
  • Khang Nguyen Hoang,
  • Khac Huy Hynh,
  • Hoang Ngoc Tran

摘要

This study evaluates models on the 30vnfood dataset to optimize Vietnamese food captioning. We propose a pipeline combining CLIP and MBart, achieving high BLEU (0.9313), CIDEr (7.4063), and METEOR (0.9311) scores. Integrating vision and language features proved essential for generating accurate captions. Our results provide actionable insights for food e-commerce and VNLP by addressing the challenges of a low-resource language in image captioning tasks. The evaluation of model performance on the Vietnamese food dataset, 30vnfood, is a critical task in determining optimal models for specific objectives. In this study, we propose a novel pipeline combining CLIP-based image embeddings with a trainable MBart model for Vietnamese caption generation. The pipeline utilizes a frozen CLIP model for feature extraction, a linear projection layer for dimensionality alignment, and MBart for natural language generation. This research systematically evaluates the effectiveness of the proposed architecture in generating captions for Vietnamese food images. Our results provide insights into the model’s performance, highlighting its potential for broader applications in food-related industries and automated systems.