Knowledge Graph-Enhanced Vision-to-Language Multimodal Models for Radiology Report Generation
摘要
Current deep learning models for automated radiology report generation leverage architecture that comprises a visual encoder and a text decoder, but often lack the semantic depth and contextual understanding necessary for producing clinically relevant, easy-to-read, and accurate reports. The situation is even more challenging due to the complex nature of medical imaging and the specialized language and medical terminologies in radiology reports. The gap in domain-specific knowledge in current deep learning models underscores the necessity for approaches that integrate specialized radiological expertise into advanced language models. In this research, we propose a knowledge graph-enhanced vision-to-language multimodal model for radiology report generation, that leverages existing medical and radiological knowledge graphs. We explore contrastive learning approaches for pre-training multimodal models to learn the joint embeddings of modalities including images, graphs and texts. Our research not only contributes to the field of semantic web research by demonstrating the potential of knowledge graphs in enhancing deep learning models but also aims to revolutionize the radiology reporting process by automating it with greater accuracy, thereby reducing the workload of radiologists and mitigating the risk of human error.