Recently, large vision-language models (VLMs) have demonstrated remarkable performance in multimodal tasks. However, their effectiveness in handling complex medical imaging remains limited. In this paper, we introduce CorGPT, an innovative multimodal medical dialogue model designed to analyze coronary angiography images and respond to open-ended questions. To bridge the gap between visual and language representations, we develop a bridge module that embeds visual features into the semantic space of a fine-tuned large language model (LLM). Additionally, we integrate a vector database (VecDB)-based autonomous retrieval mechanism to mitigate common challenges in medical AI, such as knowledge obsolescence and hallucinations. Experimental evaluations show that CorGPT outperforms baseline models, achieving significant improvements in ROUGE scores and higher consistency scores in GPT-based assessments. These results highlight its specialized capability and reliability in generating diagnostic summaries from coronary angiography images and supporting multi-turn medical dialogues. By fine-tuning on high-quality datasets, including but not limited to approximately 21,800 real-world cardiology patient-provider interaction transcripts, CorGPT achieves deep alignment between coronary angiography images and medical text, significantly enhancing diagnostic accuracy and multi-turn dialogue capability. This study provides new insights into the application of AI in medical imaging analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CorGPT: Coronary Angiography Imaging Analysis Using Large Medical Vision-Language Models

  • Baoqian Huang,
  • Hongkuan Zhang,
  • Zhaoyang Liu,
  • Jiwei Wang,
  • Shuwang Zhou

摘要

Recently, large vision-language models (VLMs) have demonstrated remarkable performance in multimodal tasks. However, their effectiveness in handling complex medical imaging remains limited. In this paper, we introduce CorGPT, an innovative multimodal medical dialogue model designed to analyze coronary angiography images and respond to open-ended questions. To bridge the gap between visual and language representations, we develop a bridge module that embeds visual features into the semantic space of a fine-tuned large language model (LLM). Additionally, we integrate a vector database (VecDB)-based autonomous retrieval mechanism to mitigate common challenges in medical AI, such as knowledge obsolescence and hallucinations. Experimental evaluations show that CorGPT outperforms baseline models, achieving significant improvements in ROUGE scores and higher consistency scores in GPT-based assessments. These results highlight its specialized capability and reliability in generating diagnostic summaries from coronary angiography images and supporting multi-turn medical dialogues. By fine-tuning on high-quality datasets, including but not limited to approximately 21,800 real-world cardiology patient-provider interaction transcripts, CorGPT achieves deep alignment between coronary angiography images and medical text, significantly enhancing diagnostic accuracy and multi-turn dialogue capability. This study provides new insights into the application of AI in medical imaging analysis.