CorGPT: Coronary Angiography Imaging Analysis Using Large Medical Vision-Language Models
摘要
Recently, large vision-language models (VLMs) have demonstrated remarkable performance in multimodal tasks. However, their effectiveness in handling complex medical imaging remains limited. In this paper, we introduce CorGPT, an innovative multimodal medical dialogue model designed to analyze coronary angiography images and respond to open-ended questions. To bridge the gap between visual and language representations, we develop a bridge module that embeds visual features into the semantic space of a fine-tuned large language model (LLM). Additionally, we integrate a vector database (VecDB)-based autonomous retrieval mechanism to mitigate common challenges in medical AI, such as knowledge obsolescence and hallucinations. Experimental evaluations show that CorGPT outperforms baseline models, achieving significant improvements in ROUGE scores and higher consistency scores in GPT-based assessments. These results highlight its specialized capability and reliability in generating diagnostic summaries from coronary angiography images and supporting multi-turn medical dialogues. By fine-tuning on high-quality datasets, including but not limited to approximately 21,800 real-world cardiology patient-provider interaction transcripts, CorGPT achieves deep alignment between coronary angiography images and medical text, significantly enhancing diagnostic accuracy and multi-turn dialogue capability. This study provides new insights into the application of AI in medical imaging analysis.