Bootstrapping multimodal large language model with medical knowledge for automatic esophagogastroduodenoscopy diagnosis and reporting
摘要
Esophagogastroduodenoscopy (EGD) is an essential clinical procedure for diagnosing gastrointestinal (GI) diseases, while the subsequent EGD reports play a vital role in clinical decision-making and therapeutic interventions. This paper makes the first attempt to bootstrap Multimodal Large Language Models (MLLM) with medical knowledge for EGD diagnosis and reporting (EDR). We collected the largest multicentric EGD dataset so far, containing 4461 participants with 203,838 EGD images and 4461 corresponding EGD reports. Experimental results demonstrate that the proposed method, MLLM-EDR, achieves an average diagnostic accuracy of 0.882 for nineteen GI diseases, outperforming both state-of-the-art AI models (0.720) and junior endoscopists (0.784). The completeness and facticity scores of our generated reports match the quality of those created by senior endoscopists. Moreover, our method substantially reduces the endoscopists’ workload from 7 minutes to 13.48 seconds. These results highlight the vigor and significance of MLLM-EDR for AI-assisted EGD diagnostic and reporting applications.