Artificial Intelligence Driven Diagnosis and Prognosis Comparison of ChatGPT-4o and DeepSeek-R1 in HIV Negative Talaromycosis
摘要
This study evaluates and compares the diagnostic and prognostic capabilities of ChatGPT-4o and DeepSeek-R1 in 56 HIV-negative talaromycosis cases. Clinical case fragments were de-identified and submitted to both models, with diagnostic accuracy and prognostic prediction rates statistically analyzed using chi-square tests, Fisher’s exact tests, and logistic regression. Results showed DeepSeek-R1 achieved significantly higher diagnostic accuracy (66.1%) than ChatGPT-4o (3.6%) (χ2 = 48.2, p < 0.001), attributable to its regional data training focusing on Southeast Asia and southern China. Conversely, ChatGPT-4o demonstrated superior prognostic prediction accuracy (78.6% vs. 50.0%, p < 0.001), with 90.2% specificity for improved (survival) outcomes, while DeepSeek-R1 showed 86.7% sensitivity for mortality. Key diagnostic predictors included hilar lymphadenectasis (odds ratio [OR] = 6.8, 95% confidence interval [CI]: 2.1–22.3, P = 0.002) and chest pain (OR = 5.9, 95% CI: 1.4–25.6, P = 0.016). The findings highlight DeepSeek-R1’s regional diagnostic advantage and ChatGPT-4o’s prognostic utility, advocating for their collaborative use to enhance early detection and management of this neglected fungal infection in immunocompromised, non-HIV populations.