Medical Visual Question Answering: A Systematic Review
摘要
This paper presents a systematic review of the medical visual question answering problem from outstanding published research in consulting and supporting medical diagnosis. The authors have constructed a structured paradigm to solve the MedVQA problem. From there, the authors focus on exploiting and analyzing the models that have been implemented. Model selection is based on a combination of image encoding (through image encoder), question encoding (through text encoder) and synthesis model (or fusion encoder). The methods are reviewed and divided into Model-based and Learning-based for image encoders. For text encoders, the authors review RNN-based approaches and Transformer-based approaches. Fusion Encoders are classified into Baseline methods, Bilinear Pooling methods and Attention mechanisms. In addition, a synthesis of Medical Vision-Language, Dataset Pre-training and Prediction module and evaluation metrics is also presented. The article provides a general picture of approaching the medical visual question answering problem most efficiently as a basis for proposing improved algorithms and methods, effectively contributing technology to the medical field.