This paper presents a systematic review of the medical visual question answering problem from outstanding published research in consulting and supporting medical diagnosis. The authors have constructed a structured paradigm to solve the MedVQA problem. From there, the authors focus on exploiting and analyzing the models that have been implemented. Model selection is based on a combination of image encoding (through image encoder), question encoding (through text encoder) and synthesis model (or fusion encoder). The methods are reviewed and divided into Model-based and Learning-based for image encoders. For text encoders, the authors review RNN-based approaches and Transformer-based approaches. Fusion Encoders are classified into Baseline methods, Bilinear Pooling methods and Attention mechanisms. In addition, a synthesis of Medical Vision-Language, Dataset Pre-training and Prediction module and evaluation metrics is also presented. The article provides a general picture of approaching the medical visual question answering problem most efficiently as a basis for proposing improved algorithms and methods, effectively contributing technology to the medical field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Medical Visual Question Answering: A Systematic Review

  • Huy Pham Le,
  • Anh Quan,
  • Thien Trang Ly,
  • Thien Doanh Le,
  • Kha Tu Huynh

摘要

This paper presents a systematic review of the medical visual question answering problem from outstanding published research in consulting and supporting medical diagnosis. The authors have constructed a structured paradigm to solve the MedVQA problem. From there, the authors focus on exploiting and analyzing the models that have been implemented. Model selection is based on a combination of image encoding (through image encoder), question encoding (through text encoder) and synthesis model (or fusion encoder). The methods are reviewed and divided into Model-based and Learning-based for image encoders. For text encoders, the authors review RNN-based approaches and Transformer-based approaches. Fusion Encoders are classified into Baseline methods, Bilinear Pooling methods and Attention mechanisms. In addition, a synthesis of Medical Vision-Language, Dataset Pre-training and Prediction module and evaluation metrics is also presented. The article provides a general picture of approaching the medical visual question answering problem most efficiently as a basis for proposing improved algorithms and methods, effectively contributing technology to the medical field.