<p>The rapid evolution of large language models (LLMs) in the medical field, particularly in automating medical tasks and supporting diagnosis and treatment, has shown promising potential. However, their accuracy, comprehensiveness, and safety in managing complex cardiovascular diseases have not been systematically assessed.&#xa0;This study aims to evaluate and compare the diagnostic, therapeutic, and safety performance of two large language models—ChatGPT-4o and Kimi—in managing complex cardiovascular diseases, and to explore their potential for future clinical application.&#xa0;A total of 200 complex cardiovascular cases published in <i>JACC: Case Reports</i> between January 2020 and August 2024 were included. All cases were standardized and de-identified before being input into ChatGPT-4o and Kimi using identical prompts. Each model independently generated diagnostic, treatment, and long-term management plans. Three cardiovascular specialists independently evaluated the outputs in a blinded manner, scoring accuracy and comprehensiveness using Likert scales. Safety was assessed using a risk matrix analysis. Additionally, 50 cases were randomly selected for triangulation to compare model-generated recommendations with clinical guidelines. Statistical analysis was performed using the Wilcoxon signed-rank test with Benjamini-Hochberg correction for multiple comparisons.&#xa0;In preliminary diagnostic accuracy, the two models performed similarly (<i>P</i> = 0.663, <i>r</i> = 0.044), but ChatGPT-4o showed superior comprehensiveness (<i>P</i> &lt; 0.001, <i>r</i> = 0.484). For treatment recommendations, ChatGPT-4o outperformed Kimi in both accuracy (<i>P</i> = 0.004, <i>r</i> = 0.321) and comprehensiveness (<i>P</i> &lt; 0.001, <i>r</i> = 0.644). In long-term management, ChatGPT-4o demonstrated significant advantages in accuracy (<i>P</i> &lt; 0.001, <i>r</i> = 0.717) and comprehensiveness (<i>P</i> &lt; 0.001, <i>r</i> = 0.690). Safety assessment showed a lower proportion of high-risk outputs with ChatGPT-4o (1.5%) compared to Kimi (4.5%).&#xa0;LLMs, particularly ChatGPT-4o, exhibit significant promise in the diagnosis and treatment of complex cardiovascular diseases, showing superior accuracy, comprehensiveness, and safety compared to Kimi. Despite their high accuracy and safety, LLMs still require clinician oversight, especially in the formulation of personalized treatment plans and complex decision-making scenarios, to ensure their reliable integration into clinical practice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Novel Insights into the Application of Large Language Models in the Diagnosis and Treatment of Complex Cardiovascular Diseases: A Comparative Study

  • Menglin Tian,
  • Shaolong Li,
  • Wenyin Du,
  • Sen Yang,
  • Xiaohua Zhao,
  • Hao Xiong,
  • Hongxi Li,
  • Mei Lu,
  • Yunyan Ying,
  • Jilei Zhang,
  • Qiwei Liao,
  • Dong Yang,
  • Fuding Guo

摘要

The rapid evolution of large language models (LLMs) in the medical field, particularly in automating medical tasks and supporting diagnosis and treatment, has shown promising potential. However, their accuracy, comprehensiveness, and safety in managing complex cardiovascular diseases have not been systematically assessed. This study aims to evaluate and compare the diagnostic, therapeutic, and safety performance of two large language models—ChatGPT-4o and Kimi—in managing complex cardiovascular diseases, and to explore their potential for future clinical application. A total of 200 complex cardiovascular cases published in JACC: Case Reports between January 2020 and August 2024 were included. All cases were standardized and de-identified before being input into ChatGPT-4o and Kimi using identical prompts. Each model independently generated diagnostic, treatment, and long-term management plans. Three cardiovascular specialists independently evaluated the outputs in a blinded manner, scoring accuracy and comprehensiveness using Likert scales. Safety was assessed using a risk matrix analysis. Additionally, 50 cases were randomly selected for triangulation to compare model-generated recommendations with clinical guidelines. Statistical analysis was performed using the Wilcoxon signed-rank test with Benjamini-Hochberg correction for multiple comparisons. In preliminary diagnostic accuracy, the two models performed similarly (P = 0.663, r = 0.044), but ChatGPT-4o showed superior comprehensiveness (P < 0.001, r = 0.484). For treatment recommendations, ChatGPT-4o outperformed Kimi in both accuracy (P = 0.004, r = 0.321) and comprehensiveness (P < 0.001, r = 0.644). In long-term management, ChatGPT-4o demonstrated significant advantages in accuracy (P < 0.001, r = 0.717) and comprehensiveness (P < 0.001, r = 0.690). Safety assessment showed a lower proportion of high-risk outputs with ChatGPT-4o (1.5%) compared to Kimi (4.5%). LLMs, particularly ChatGPT-4o, exhibit significant promise in the diagnosis and treatment of complex cardiovascular diseases, showing superior accuracy, comprehensiveness, and safety compared to Kimi. Despite their high accuracy and safety, LLMs still require clinician oversight, especially in the formulation of personalized treatment plans and complex decision-making scenarios, to ensure their reliable integration into clinical practice.