Medical Visual Question Answering (VQA) has become increasingly significant in aiding physicians with disease diagnosis and providing patients with detailed insights into their conditions. However, Medical VQA still lags behind general VQA due to challenges such as limited availability of accurate data and the complexity of medical terminology. Existing models frequently struggle with performance issues caused by the intricate nature of image and text encoders. Recent research has concentrated on improving fusion modules for integrating question and image features and employing pre-trained models with self-collected datasets, but often overlooks the value of question and image history. This paper proposes an approach that introduces an Associative Memory block to leverage historical questions and images, enhancing the vision-language context. Additionally, we incorporate a Prototype Learning block that utilizes hierarchical prototype learning on text and image embeddings through advanced Hopfield layers. Our approach focuses on identifying the most representative prototypes from text-image embeddings, enriched by associative memory, rather than learning direct text-image joint feature representations. This method facilitates a more nuanced representation of semantics for answering questions. Our proposed approach achieves the best performance on the VQA-RAD dataset, demonstrating a significant accuracy improvement of 0.45%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual Question Answering for Medical Data Using a Visio-Linguistic Model

  • Van Hieu Bui,
  • Quang Duc Tran,
  • Kim Thanh Tran Thi,
  • Viet Tien Le

摘要

Medical Visual Question Answering (VQA) has become increasingly significant in aiding physicians with disease diagnosis and providing patients with detailed insights into their conditions. However, Medical VQA still lags behind general VQA due to challenges such as limited availability of accurate data and the complexity of medical terminology. Existing models frequently struggle with performance issues caused by the intricate nature of image and text encoders. Recent research has concentrated on improving fusion modules for integrating question and image features and employing pre-trained models with self-collected datasets, but often overlooks the value of question and image history. This paper proposes an approach that introduces an Associative Memory block to leverage historical questions and images, enhancing the vision-language context. Additionally, we incorporate a Prototype Learning block that utilizes hierarchical prototype learning on text and image embeddings through advanced Hopfield layers. Our approach focuses on identifying the most representative prototypes from text-image embeddings, enriched by associative memory, rather than learning direct text-image joint feature representations. This method facilitates a more nuanced representation of semantics for answering questions. Our proposed approach achieves the best performance on the VQA-RAD dataset, demonstrating a significant accuracy improvement of 0.45%.