In this paper, we explore the Bootstrapping Language Image Pretraining (BLIP) model’s performance in Visual Question Answering (VQA) tasks, particularly in the medical domain. We aim to propose an effective approach for medical image analysis by fine-tuning and optimizing the BLIP model using the Low-Rank Adaptation (LoRA) adaptation. The optimized model is tested on a combination of benchmark datasets such as MED-2019, VQA-RAD, and SLAKE-English for variety and randomness. We obtained an overall test accuracy of 75.67% for VQA-RAD and 80.9% for MED-2019 and SLAKE-ENGLISH, which highlights the potential of the LoRA-enhanced BLIP model in promoting healthcare solutions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Vision Language Model for Visual Question Answering in Medical Images

  • M. R. Dinesh Kumar,
  • Pillalamarri Akshaya,
  • R. Saivarsha,
  • N. T. Shrish Surya,
  • B. Premjith,
  • V. Sowmya,
  • G. Jyothish Lal

摘要

In this paper, we explore the Bootstrapping Language Image Pretraining (BLIP) model’s performance in Visual Question Answering (VQA) tasks, particularly in the medical domain. We aim to propose an effective approach for medical image analysis by fine-tuning and optimizing the BLIP model using the Low-Rank Adaptation (LoRA) adaptation. The optimized model is tested on a combination of benchmark datasets such as MED-2019, VQA-RAD, and SLAKE-English for variety and randomness. We obtained an overall test accuracy of 75.67% for VQA-RAD and 80.9% for MED-2019 and SLAKE-ENGLISH, which highlights the potential of the LoRA-enhanced BLIP model in promoting healthcare solutions.