The Medical Imaging Question Answering task combines medical imaging and natural language processing to answer questions related to medical imaging. Despite the progress that has been made in the field, problems remain. Currently, most image encoders use a transformer structure to extract features and output the final layer of the model for further processing. This approach ignores the complex semantic context between image and text, limiting the ability of the model to capture cross-modal semantics. To address this limitation and further explore the semantic interactions between images and text, this paper designs a Contextual Interactive Attention Connection module. The module utilises deep and shallow feature representations of the encoder and applies a variant of the attention mechanism to enable deep interaction between image and text features. This greatly improves semantic consistency and overall performance in medical vision question answering tasks. Considering that accurate answers to specialised medical questions often depend on rich medical a priori knowledge, effective integration of this knowledge to improve the accuracy of a question answering system is very costly in terms of human and financial resources. To address this problem, this paper proposes a learning matrix assistance module that utilises a learning matrix to assist the model. Experiments on two datasets, VQA-RAD and SLAKE, show that the model proposed in this paper outperforms other state-of-the-art models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Contextual Feature-Based Medical Visual Question Answering Aided by Learnable Matrix

  • Cheng Gong,
  • Haiwei Pan,
  • Haiyan Lan,
  • Kejia Zhang,
  • Shuning He,
  • Xiteng Jia

摘要

The Medical Imaging Question Answering task combines medical imaging and natural language processing to answer questions related to medical imaging. Despite the progress that has been made in the field, problems remain. Currently, most image encoders use a transformer structure to extract features and output the final layer of the model for further processing. This approach ignores the complex semantic context between image and text, limiting the ability of the model to capture cross-modal semantics. To address this limitation and further explore the semantic interactions between images and text, this paper designs a Contextual Interactive Attention Connection module. The module utilises deep and shallow feature representations of the encoder and applies a variant of the attention mechanism to enable deep interaction between image and text features. This greatly improves semantic consistency and overall performance in medical vision question answering tasks. Considering that accurate answers to specialised medical questions often depend on rich medical a priori knowledge, effective integration of this knowledge to improve the accuracy of a question answering system is very costly in terms of human and financial resources. To address this problem, this paper proposes a learning matrix assistance module that utilises a learning matrix to assist the model. Experiments on two datasets, VQA-RAD and SLAKE, show that the model proposed in this paper outperforms other state-of-the-art models.