Medical Visual Question Answering (Med-VQA) targets at answering a clinical question associated with a corresponding medical image. Med-VQA has revealed huge potential in the field of medicine, but it still faces many challenges in practice. Existing Med-VQA models have not fully utilized medical answer features, though Compact Trilinear Interaction (CTI) model has proven that the answer has close correlations with question and image. However, directly applying existing CTI model in general domain on the small volume of Med-VQA datasets would lead to over-fitting and obtain poor performance. Therefore, this paper proposes a novel trilinear distillation learning framework called TDL for Med-VQA to learn correlations between medical answer, image and question from the trilinear model by distillation learning. In addition, considering the clinical questions are harder to understand due to the professionalism of medical scenarios, we design a question feature capturing (QFC) module to capture the fine-grained intra-modality relationships and characteristics of clinical questions. Furthermore, we take account of the unbalanced-labels of Med-VQA datasets, and propose a novel Label Smoothing Regularization Focal Loss to enhance the generalization capability of the model while dynamically adjust the sample weights during training. Based on the standard benchmark dataset VQA-RAD, our proposed TDL model achieves the state-of-the-art performance, including the best overall Accuracy 76.71%, Precision 83.78%, Recall 76.71% and F1-score 78.94%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Trilinear Distillation Learning and Question Feature Capturing for Medical Visual Question Answering

  • Shaopei Long,
  • Yong Li,
  • Heng Weng,
  • Buzhou Tang,
  • Fu Lee Wang,
  • Tianyong Hao

摘要

Medical Visual Question Answering (Med-VQA) targets at answering a clinical question associated with a corresponding medical image. Med-VQA has revealed huge potential in the field of medicine, but it still faces many challenges in practice. Existing Med-VQA models have not fully utilized medical answer features, though Compact Trilinear Interaction (CTI) model has proven that the answer has close correlations with question and image. However, directly applying existing CTI model in general domain on the small volume of Med-VQA datasets would lead to over-fitting and obtain poor performance. Therefore, this paper proposes a novel trilinear distillation learning framework called TDL for Med-VQA to learn correlations between medical answer, image and question from the trilinear model by distillation learning. In addition, considering the clinical questions are harder to understand due to the professionalism of medical scenarios, we design a question feature capturing (QFC) module to capture the fine-grained intra-modality relationships and characteristics of clinical questions. Furthermore, we take account of the unbalanced-labels of Med-VQA datasets, and propose a novel Label Smoothing Regularization Focal Loss to enhance the generalization capability of the model while dynamically adjust the sample weights during training. Based on the standard benchmark dataset VQA-RAD, our proposed TDL model achieves the state-of-the-art performance, including the best overall Accuracy 76.71%, Precision 83.78%, Recall 76.71% and F1-score 78.94%.