Automatic radiology report generation with deep learning: a comprehensive review of methods and advances
摘要
Automatic report generation refers to the process of generating medical reports from medical images without the need for manual intervention, enabling faster, more consistent, and objective analysis of radiological data. The rapid progress in deep learning, particularly in the fields of computer vision and natural language processing, has significantly improved the efficacy of this approach. By leveraging deep learning techniques, which seamlessly integrate image analysis with natural language generation, these methods have shown promise in interpreting complex medical images and producing highly accurate textual descriptions. In this paper, we provide a thorough review of various deep learning models and techniques employed for generating radiological reports, with a focus on chest X-ray images as a representative case. We propose a unified encoder-decoder framework that consists of an image encoder for extracting feature representations from medical images, a language decoder for generating textual reports, and enhancement components designed to refine model performance. Through a comprehensive comparison of existing state-of-the-art methods on the widely utilized MIMIC-CXR dataset, we highlight the innovative contributions made by recent advancements in the field. Furthermore, we discuss the current challenges and identify potential research directions for future advancements in this field.