Image Captioning is a process where a computer automatically generates textual descriptions to interpret the content of an image. Scene Graphs, as a hierarchical method of organizing data, can label the semantic connections between objects in an image. By creating a scene graph for an image, these relationships can be integrated into the image captioning model, thereby enhancing the capture of local features and aiding in the derivation of accurate textual descriptions. However, existing scene graph generation techniques often produce a large amount of unnecessary redundant information. To fully exploit the semantic information in scene graphs that is beneficial for description and reduce the interference of redundant information, this paper proposes an encoder that combines graph attention mechanisms and graph update mechanisms after constructing the image's scene graph. This encoder can automatically identify and focus on key information that contributes to description generation, aggregating these relationships to form locally feature-aware representations. Furthermore, to improve the model's attention efficiency and accuracy in focusing on the image content during the description generation process, this paper introduces an adaptive attention module based on the visual sentinel mechanism at the decoding stage of description generation. This module can more flexibly switch between relying on visual signals and relying on language models. Experimental validation on the MS-COCO image captioning dataset has shown that the method proposed in this paper surpasses all the latest methods that utilize semantic relationships to guide image captioning in performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Captioning Based on Scene Graph Optimization

  • Zi Wang,
  • He Wang,
  • Huazhou Hou,
  • Wenwu Yu

摘要

Image Captioning is a process where a computer automatically generates textual descriptions to interpret the content of an image. Scene Graphs, as a hierarchical method of organizing data, can label the semantic connections between objects in an image. By creating a scene graph for an image, these relationships can be integrated into the image captioning model, thereby enhancing the capture of local features and aiding in the derivation of accurate textual descriptions. However, existing scene graph generation techniques often produce a large amount of unnecessary redundant information. To fully exploit the semantic information in scene graphs that is beneficial for description and reduce the interference of redundant information, this paper proposes an encoder that combines graph attention mechanisms and graph update mechanisms after constructing the image's scene graph. This encoder can automatically identify and focus on key information that contributes to description generation, aggregating these relationships to form locally feature-aware representations. Furthermore, to improve the model's attention efficiency and accuracy in focusing on the image content during the description generation process, this paper introduces an adaptive attention module based on the visual sentinel mechanism at the decoding stage of description generation. This module can more flexibly switch between relying on visual signals and relying on language models. Experimental validation on the MS-COCO image captioning dataset has shown that the method proposed in this paper surpasses all the latest methods that utilize semantic relationships to guide image captioning in performance.