<p>Symptom-Assessment Application (SAAs, e.g., NHS 111 online) that assist laypeople in deciding if and where to seek care (<i>self-triage</i>) are gaining popularity and Large Language Models (LLMs) are increasingly used too. However, there is no evidence synthesis on the accuracy of LLMs, and no review has contextualized the accuracy of SAAs and LLMs. This systematic review evaluates the self-triage accuracy of both SAAs and LLMs and compares them to the accuracy of laypeople. A total of 1549 studies were screened and 19 included. The self-triage accuracy of SAAs was moderate but highly variable (11.5–90.0%), while the accuracy of LLMs (57.8–76.0%) and laypeople (47.3–62.4%) was moderate with low variability. Based on the available evidence, the use of SAAs or LLMs should neither be universally recommended nor discouraged; rather, we suggest that their utility should be assessed based on the specific use case and user group under consideration.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accuracy of online symptom assessment applications, large language models, and laypeople for self–triage decisions

  • Marvin Kopka,
  • Niklas von Kalckreuth,
  • Markus A. Feufel

摘要

Symptom-Assessment Application (SAAs, e.g., NHS 111 online) that assist laypeople in deciding if and where to seek care (self-triage) are gaining popularity and Large Language Models (LLMs) are increasingly used too. However, there is no evidence synthesis on the accuracy of LLMs, and no review has contextualized the accuracy of SAAs and LLMs. This systematic review evaluates the self-triage accuracy of both SAAs and LLMs and compares them to the accuracy of laypeople. A total of 1549 studies were screened and 19 included. The self-triage accuracy of SAAs was moderate but highly variable (11.5–90.0%), while the accuracy of LLMs (57.8–76.0%) and laypeople (47.3–62.4%) was moderate with low variability. Based on the available evidence, the use of SAAs or LLMs should neither be universally recommended nor discouraged; rather, we suggest that their utility should be assessed based on the specific use case and user group under consideration.