Brain imaging reports often contain complex medical jargon that is difficult for patients without a medical background to understand. In this study, we examined the ability of four popular large language models (LLMs; ChatGPT-3.5, ChatGPT-4, Google Bard, and Microsoft Bing) to simplify brain magnetic resonance imaging (MRI) reports for patients and compared physicians’ and patients’ satisfaction with the simplified reports generated by various types of LLMs. Specifically, 24 physicians and 40 participants without a medical background evaluated simplified reports generated from a full MRI report by each of the four LLMs, respectively. The results showed that physicians were satisfied with the comprehensibility, factual correctness, and completeness of the four versions of the simplified reports, but were neutral regarding the potential harm and overall quality of the simplified reports. They were also less likely to send the simplified reports to patients. Additionally, physicians were more satisfied with the reports generated by Microsoft Bing and ChatGPT-4 compared to those generated by the other two LLMs, in terms of potential harm, overall quality, and likelihood of sending the report to patients. Participants, on the other hand, reported that they easily understood the content of the simplified reports, drew correct conclusions, and expressed a high willingness to receive the simplified reports. There was no significant difference in participants’ ratings of the four simplified reports. These findings suggest great potential for using AI models to improve patient-centered care in brain imaging and other medical fields.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing the Feasibility of Using AI Models to Simplify Brain Imaging Reports for Patients: A Comparative Analysis of Four Large Language Models

  • Min Xu,
  • Yiwen Wang

摘要

Brain imaging reports often contain complex medical jargon that is difficult for patients without a medical background to understand. In this study, we examined the ability of four popular large language models (LLMs; ChatGPT-3.5, ChatGPT-4, Google Bard, and Microsoft Bing) to simplify brain magnetic resonance imaging (MRI) reports for patients and compared physicians’ and patients’ satisfaction with the simplified reports generated by various types of LLMs. Specifically, 24 physicians and 40 participants without a medical background evaluated simplified reports generated from a full MRI report by each of the four LLMs, respectively. The results showed that physicians were satisfied with the comprehensibility, factual correctness, and completeness of the four versions of the simplified reports, but were neutral regarding the potential harm and overall quality of the simplified reports. They were also less likely to send the simplified reports to patients. Additionally, physicians were more satisfied with the reports generated by Microsoft Bing and ChatGPT-4 compared to those generated by the other two LLMs, in terms of potential harm, overall quality, and likelihood of sending the report to patients. Participants, on the other hand, reported that they easily understood the content of the simplified reports, drew correct conclusions, and expressed a high willingness to receive the simplified reports. There was no significant difference in participants’ ratings of the four simplified reports. These findings suggest great potential for using AI models to improve patient-centered care in brain imaging and other medical fields.