Background <p>ChatGPT has demonstrated strong performance in the complex, full clinical workflow. In recent years, several large language models (LLMs) from China have been introduced; however, their performance in such intricate tasks has yet to be thoroughly assessed, and it remains unclear whether their performance diverges from that of ChatGPT. This study seeks to evaluate the capacity of the Chinese LLMs for providing continuous clinical decision support by assessing their performance with simulated patient cases.</p> Methods <p>We selected 29 standard cases from the <i>Merck Manual</i> as simulated patients. We provided their information to the LLMs. Each simulated case is accompanied by a series of sequential questions designed to simulate the process of differential diagnosis, diagnostic workup, diagnosis and management. The responses were then recorded and scored. Then we compared the performance of two Chinese large language models with ChatGPT-4 in entire clinical workflow of simulated patient, selecting the best-performing model for a comparison with 18 human emergency fellow doctors. Additionally, we compared the differences in performance between different versions of the LLMs.</p> Results <p>There were no significant differences between ChatGPT-4 and Doubao in all four aspects (<i>P</i> &gt; 0.05). However, ERNIE Bot 3.5 was inferior to ChatGPT-4 and Doubao in differential diagnosis, diagnostic questions, and management (<i>P</i> &lt; 0.05). But in diagnosis questions, the average accuracy proportion for all three models was above 97%, with no significant differences observed (<i>P</i> &gt; 0.05). There was no significant difference between LLMs and emergency fellow physicians in diagnosis and differential diagnosis (<i>P</i> &gt; 0.05), but in diagnostic questions as well as management, LMMs were superior to emergency fellow physicians (<i>P</i> &lt; 0.05). ChatGPT-4 was higher than ChatGPT 3.5 in all four aspects (<i>P</i> &lt; 0.05).</p> Conclusion <p>The large language model Doubao from China demonstrates performance similar to ChatGPT across full clinical workflows. LLMs outperform human emergency fellow physicians and exhibit rapid development, offering significant practical application potential in healthcare.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comparison of the performance of Chinese large language models and ChatGPT throughout the entire clinical workflow

  • Yang He,
  • Yang Wang,
  • Lingling Lai,
  • Ning Liu,
  • Yucai Hong,
  • PengPeng Chen,
  • Zhongheng Zhang

摘要

Background

ChatGPT has demonstrated strong performance in the complex, full clinical workflow. In recent years, several large language models (LLMs) from China have been introduced; however, their performance in such intricate tasks has yet to be thoroughly assessed, and it remains unclear whether their performance diverges from that of ChatGPT. This study seeks to evaluate the capacity of the Chinese LLMs for providing continuous clinical decision support by assessing their performance with simulated patient cases.

Methods

We selected 29 standard cases from the Merck Manual as simulated patients. We provided their information to the LLMs. Each simulated case is accompanied by a series of sequential questions designed to simulate the process of differential diagnosis, diagnostic workup, diagnosis and management. The responses were then recorded and scored. Then we compared the performance of two Chinese large language models with ChatGPT-4 in entire clinical workflow of simulated patient, selecting the best-performing model for a comparison with 18 human emergency fellow doctors. Additionally, we compared the differences in performance between different versions of the LLMs.

Results

There were no significant differences between ChatGPT-4 and Doubao in all four aspects (P > 0.05). However, ERNIE Bot 3.5 was inferior to ChatGPT-4 and Doubao in differential diagnosis, diagnostic questions, and management (P < 0.05). But in diagnosis questions, the average accuracy proportion for all three models was above 97%, with no significant differences observed (P > 0.05). There was no significant difference between LLMs and emergency fellow physicians in diagnosis and differential diagnosis (P > 0.05), but in diagnostic questions as well as management, LMMs were superior to emergency fellow physicians (P < 0.05). ChatGPT-4 was higher than ChatGPT 3.5 in all four aspects (P < 0.05).

Conclusion

The large language model Doubao from China demonstrates performance similar to ChatGPT across full clinical workflows. LLMs outperform human emergency fellow physicians and exhibit rapid development, offering significant practical application potential in healthcare.