Streamlined Speech Recognition Model for Automated Radiology Reporting Employing Combined Automatic Speech Recognition Model, Large Language Model, and Prompt Engineering
摘要
As radiologists around the globe contend with ever-growing workloads, the implementation of automated report generation presents a dual advantage: easing their burden and improving turnaround times. Empowered by Artificial Intelligence (AI)-driven speech recognition, our proposed method minimizes dictation errors, enhancing precision in report generation and thereby reducing the need for post-dictation corrections. In our research, we aim to pinpoint the most effective end-to-end Dictation-to-Report Pipeline. This entails assessing the performance of various combinations of NLP models (Facebook Word2vec Whisper, Facebook Word2letter, Kaldi, DeepSpeech), LLM models (Mistral-7B, Falcon-7B, Zephyr-7B, Qwen-14B), and Prompt Engineering for speech recognition and subsequent report generation. The achieved results include a Match Error Rate of 0.016, Word Error Rate of 0.131, Average Levenshtein Distance of 0.109, and Sentence Error Rate of 0.019. Our conclusions are further substantiated through a blinded qualitative evaluation by five radiologists, who assign validation scores to the generated radiology reports.