<p>Visual language pre-training (VLP) models show strong potential in learning knowledge from large-scale data, but still fall short in dealing with novel entity naming and dynamic real-world knowledge representation. Traditional methods rely on collecting a large amount of text corresponding to images for retraining, which is both time-consuming and difficult to adapt to the rapid update of knowledge. To this end, this paper proposes a Key Knowledge Prompt Enhancement (KKPE) approach, which introduces key knowledge as prompts, uses different prompts on different training samples during the training process, and enables the model to learn how to integrate key knowledge into the generated descriptions through a prompt differentiation strategy. In the inference phase, the model is able to generate knowledge-enhanced image descriptions based on the prompts, even if they are novel real-world knowledge. The approach does not require large-scale retraining and is able to respond quickly to real-world knowledge changes, which significantly improves the model’s generative power and flexibility. Experimental results show that our method is able to effectively incorporate key knowledge into descriptions through key knowledge prompts, with significant improvements in CIDEr scores and knowledge recognition accuracies on the KnowCap dataset compared to the baseline model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced image description generation based on key knowledge prompt

  • Jianwen Mo,
  • Jilin He,
  • Hua Yuan

摘要

Visual language pre-training (VLP) models show strong potential in learning knowledge from large-scale data, but still fall short in dealing with novel entity naming and dynamic real-world knowledge representation. Traditional methods rely on collecting a large amount of text corresponding to images for retraining, which is both time-consuming and difficult to adapt to the rapid update of knowledge. To this end, this paper proposes a Key Knowledge Prompt Enhancement (KKPE) approach, which introduces key knowledge as prompts, uses different prompts on different training samples during the training process, and enables the model to learn how to integrate key knowledge into the generated descriptions through a prompt differentiation strategy. In the inference phase, the model is able to generate knowledge-enhanced image descriptions based on the prompts, even if they are novel real-world knowledge. The approach does not require large-scale retraining and is able to respond quickly to real-world knowledge changes, which significantly improves the model’s generative power and flexibility. Experimental results show that our method is able to effectively incorporate key knowledge into descriptions through key knowledge prompts, with significant improvements in CIDEr scores and knowledge recognition accuracies on the KnowCap dataset compared to the baseline model.