A Comprehensive Review of Medical Image Captioning: Techniques, Challenges, and Future Directions
摘要
Medical image captioning is an emerging field that combines natural language processing and computer vision for developing textual descriptions of medical images, with a focus on recent developments in deep learning models; particularly CNN, RNN and transformer-based architectures. This review aims to discuss state-of-the-art approaches and strategies applied in medical image captioning. The different datasets used in training and testing these models are also discussed by the paper, along with the problems and challenges that exist in this field. Through the exploration of the current trends, this paper attempts to give a comprehensive review of the development of medical image captioning, identify gaps in existing research, and propose new potential areas for further work. Medical image captioning improves diagnostic accuracy and automates routine reporting, reducing the work for the radiologists. Although progress, challenges continue to arise. One important problem is the lack of large, annotated medical datasets required for efficient model construction. Accurate comprehension and creation of medical terminology is crucial to avoid inaccurate captions and misdiagnosis. As the research in the field advances so that the future works should focuses issues on extensive datasets and creating model that can trained from limited data ensuring clinically accurate captions. Integrating multiple data like patient history patient history or lab reports can improve the accuracy and value of medical image captions, pushing the boundaries of automated findings. Using Artificial Intelligence in the medical field should consider some of the concerns like patient safety and privacy. The collaboration between AI research and medical professionals is essential for developing and accurate medical image caption generator.