<p>Controllable image captioning has become a hot research topic in recent years by controlling the parameters or conditions in the generation process to obtain descriptions that meet the user’s expectations. However, the captions generated by the current controllable caption model usually contain only a single simple sentence, and the sentences tend to be templated without rich linguistic expressions, resulting in monotonous and boring descriptions. To solve these problems, this paper proposes an optimization method for controllable image caption generation based on scene graph sorting-selection (SS) and shuffle-polishing (SP). First, we utilize the scene graph technique to capture the complex object relationships in images and decompose them into multiple substructures, and then introduce a new two-dimensional matching score algorithm to sort the decomposed multiple substructures. The selected substructures are decoded into multiple target sentences based on the sorting result or user intent, which describe the image content from different perspectives to increase the diversity and richness of the generated results. Subsequently, in the shuffle polishing stage, we take the sentences initially generated by the language decoder module as complete contextual information, and utilize the rich linguistic structure as well as vocabulary knowledge learned from the unsupervised model to iteratively update the words in each position of the sentences in a randomized order, to ultimately generate flexible yet personalized descriptions. Extensive experiments on MSCOCO and Flickr30k Entities datasets validate the excellent performance of our model in controllable image captioning, generating novel and diverse captions while taking into account the accuracy of descriptions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scene graph sorting and shuffle polishing based controllable image captioning

  • Guichang Wu,
  • Qian Zhao,
  • Xiushu Liu

摘要

Controllable image captioning has become a hot research topic in recent years by controlling the parameters or conditions in the generation process to obtain descriptions that meet the user’s expectations. However, the captions generated by the current controllable caption model usually contain only a single simple sentence, and the sentences tend to be templated without rich linguistic expressions, resulting in monotonous and boring descriptions. To solve these problems, this paper proposes an optimization method for controllable image caption generation based on scene graph sorting-selection (SS) and shuffle-polishing (SP). First, we utilize the scene graph technique to capture the complex object relationships in images and decompose them into multiple substructures, and then introduce a new two-dimensional matching score algorithm to sort the decomposed multiple substructures. The selected substructures are decoded into multiple target sentences based on the sorting result or user intent, which describe the image content from different perspectives to increase the diversity and richness of the generated results. Subsequently, in the shuffle polishing stage, we take the sentences initially generated by the language decoder module as complete contextual information, and utilize the rich linguistic structure as well as vocabulary knowledge learned from the unsupervised model to iteratively update the words in each position of the sentences in a randomized order, to ultimately generate flexible yet personalized descriptions. Extensive experiments on MSCOCO and Flickr30k Entities datasets validate the excellent performance of our model in controllable image captioning, generating novel and diverse captions while taking into account the accuracy of descriptions.