<p>Most current Visual Question Answering (VQA) methods struggle to achieve effective cross-modal interaction between visual and semantic information, resulting in difficulties in accurately combining visual content with contextual semantics for answer prediction. To address this problem, a Cross-modal Heterogeneous Graph Reasoning Network (CHGRN) is proposed for VQA, incorporating a novel Cross-modal Reasoning Module (CRM) to enhance the interactive analysis between images and questions, enabling more profound joint reasoning of visual and semantic features. The CRM improves cross-modal information reasoning by effectively analyzing and understanding the visual semantic information in the target regions of images. Additionally, an answer type prediction module is introduced, employing multi-task learning with answer type annotations to filter out irrelevant semantic information, thereby improving reasoning accuracy. Moreover, the semantic-assisted attention-aligned decoder ensures precise alignment between visual and semantic data. Extensive experiments demonstrate that the proposed CHGRN achieves excellent performance in visual question answering and outperforms most state-of-the-art methods on widely used public datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-modal heterogeneous graph reasoning network for visual question answering

  • Jing Zhang,
  • Jiong Teng,
  • Weichao Ding,
  • Zhe Wang

摘要

Most current Visual Question Answering (VQA) methods struggle to achieve effective cross-modal interaction between visual and semantic information, resulting in difficulties in accurately combining visual content with contextual semantics for answer prediction. To address this problem, a Cross-modal Heterogeneous Graph Reasoning Network (CHGRN) is proposed for VQA, incorporating a novel Cross-modal Reasoning Module (CRM) to enhance the interactive analysis between images and questions, enabling more profound joint reasoning of visual and semantic features. The CRM improves cross-modal information reasoning by effectively analyzing and understanding the visual semantic information in the target regions of images. Additionally, an answer type prediction module is introduced, employing multi-task learning with answer type annotations to filter out irrelevant semantic information, thereby improving reasoning accuracy. Moreover, the semantic-assisted attention-aligned decoder ensures precise alignment between visual and semantic data. Extensive experiments demonstrate that the proposed CHGRN achieves excellent performance in visual question answering and outperforms most state-of-the-art methods on widely used public datasets.