<p>This study evaluates and compares the diagnostic and prognostic capabilities of ChatGPT-4o and DeepSeek-R1 in 56 HIV-negative talaromycosis cases. Clinical case fragments were de-identified and submitted to both models, with diagnostic accuracy and prognostic prediction rates statistically analyzed using chi-square tests, Fisher’s exact tests, and logistic regression. Results showed DeepSeek-R1 achieved significantly higher diagnostic accuracy (66.1%) than ChatGPT-4o (3.6%) (χ<sup>2</sup> = 48.2, <i>p</i> &lt; 0.001), attributable to its regional data training focusing on Southeast Asia and southern China. Conversely, ChatGPT-4o demonstrated superior prognostic prediction accuracy (78.6% vs. 50.0%, <i>p</i> &lt; 0.001), with 90.2% specificity for improved (survival) outcomes, while DeepSeek-R1 showed 86.7% sensitivity for mortality. Key diagnostic predictors included hilar lymphadenectasis (odds ratio [<i>OR</i>] = 6.8, 95% confidence interval [<i>CI</i>]: 2.1–22.3,&#xa0;<i>P</i> = 0.002) and chest pain (<i>OR</i> = 5.9, 95% <i>CI</i>: 1.4–25.6,&#xa0;<i>P</i> = 0.016). The findings highlight DeepSeek-R1’s regional diagnostic advantage and ChatGPT-4o’s prognostic utility, advocating for their collaborative use to enhance early detection and management of this neglected fungal infection in immunocompromised, non-HIV populations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial Intelligence Driven Diagnosis and Prognosis Comparison of ChatGPT-4o and DeepSeek-R1 in HIV Negative Talaromycosis

  • Haiyang He,
  • Liuyang Cai,
  • Yi Liu,
  • Yusong Lin,
  • Xingrui Zhu,
  • Dongzhen Liu,
  • Wanqing Liao,
  • Xiaochun Xue,
  • Weihua Pan

摘要

This study evaluates and compares the diagnostic and prognostic capabilities of ChatGPT-4o and DeepSeek-R1 in 56 HIV-negative talaromycosis cases. Clinical case fragments were de-identified and submitted to both models, with diagnostic accuracy and prognostic prediction rates statistically analyzed using chi-square tests, Fisher’s exact tests, and logistic regression. Results showed DeepSeek-R1 achieved significantly higher diagnostic accuracy (66.1%) than ChatGPT-4o (3.6%) (χ2 = 48.2, p < 0.001), attributable to its regional data training focusing on Southeast Asia and southern China. Conversely, ChatGPT-4o demonstrated superior prognostic prediction accuracy (78.6% vs. 50.0%, p < 0.001), with 90.2% specificity for improved (survival) outcomes, while DeepSeek-R1 showed 86.7% sensitivity for mortality. Key diagnostic predictors included hilar lymphadenectasis (odds ratio [OR] = 6.8, 95% confidence interval [CI]: 2.1–22.3, P = 0.002) and chest pain (OR = 5.9, 95% CI: 1.4–25.6, P = 0.016). The findings highlight DeepSeek-R1’s regional diagnostic advantage and ChatGPT-4o’s prognostic utility, advocating for their collaborative use to enhance early detection and management of this neglected fungal infection in immunocompromised, non-HIV populations.