Visual question generation task aims to generate meaningful questions about an image targeting an answer. Despite we made significant progress in automatically generating questions on the characteristics of visual objects contained in images, existing methods often ignore text content related to visual objects in images, which helps people understand the image and perform some daily tasks. To address this problem, we propose a texts-aware generation model that extracts the text content and visual objects related to the answer to generate questions. Considering that the relationships among visual objects and among text content are helpful for this task, we propose a dual GCN module to extract object-level and word-level representations of images containing these relationships. Each representation contributes differently to a generated question, so we try to extract image regions that are relevant to an answer for questioning. To extract the visual object related to the answer, we introduce a correlation calculation module to calculate the relationship between each visual object and the answer based on the positional relationships. To capture the text content related to the answer, our model utilizes an answer-aware module to calculate the relationship between each word and the answer. Finally, we mainly focus on the answer-related text content and visual object features in an image to generate questions. Extensive experiments on the TextVQA dataset and the ST-VQA dataset show that the proposed model outperforms the existing models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning to Ask About Text Content in an Image with Fine-Grained Features

  • Mengqiu Cheng,
  • Wanting Qiao,
  • Junming Chen,
  • Xianfeng Li

摘要

Visual question generation task aims to generate meaningful questions about an image targeting an answer. Despite we made significant progress in automatically generating questions on the characteristics of visual objects contained in images, existing methods often ignore text content related to visual objects in images, which helps people understand the image and perform some daily tasks. To address this problem, we propose a texts-aware generation model that extracts the text content and visual objects related to the answer to generate questions. Considering that the relationships among visual objects and among text content are helpful for this task, we propose a dual GCN module to extract object-level and word-level representations of images containing these relationships. Each representation contributes differently to a generated question, so we try to extract image regions that are relevant to an answer for questioning. To extract the visual object related to the answer, we introduce a correlation calculation module to calculate the relationship between each visual object and the answer based on the positional relationships. To capture the text content related to the answer, our model utilizes an answer-aware module to calculate the relationship between each word and the answer. Finally, we mainly focus on the answer-related text content and visual object features in an image to generate questions. Extensive experiments on the TextVQA dataset and the ST-VQA dataset show that the proposed model outperforms the existing models.