Assessment of ChatGPT’s adherence to EULAR diagnostic criteria and therapeutic protocols for rheumatoid arthritis at two distinct time points, 14 days apart, utilizing binary and multiple-choice inquiries
摘要
Artificial intelligence (AI) possesses considerable promise in healthcare to offer decision help in particular domains, including rheumatoid arthritis (RA). This study assesses the adherence of the advanced AI model ChatGPT-v4 to the European League Against Rheumatism (EULAR) recommendations.
MethodsThe research employed a 100-item questionnaire consisting of true/false and multiple-choice formats, accompanied with real-world clinical scenarios developed concurrently with EULAR in the therapy of RA. Inquiries addressed diagnostic criteria, therapeutic alternatives, and follow-up procedures. Two rheumatologists assessed the ChatGPT for accuracy, consistency, and comprehensiveness utilizing a 6-point Likert scale.
ResultsEvaluation occurred at baseline and on day 14. AI rectified the majority of errors at baseline in the paired questions. It did not advance on specific responses. One of the two previously incongruent responses remained unaltered, while the other was rectified. The 48 originally congruent responses rose to 49 on day 14. In binary questions, AI exhibited greater coherence than in multiple-choice questions. At baseline, 43 (86%) of the multiple-choice items were answered correctly. Upon reevaluation, 42 (84%) were found to be accurate. One response was erroneous on day 14. Three of the seven initially erroneous responses remained unaltered. Four erroneous responses were later rectified.
ConclusionChatGPT demonstrated efficacy in addressing binary and multiple-choice questions formulated according to EULAR guidelines for RA. The findings validated that AI can serve as a clinical support instrument in RA. It demonstrated that AI can be enhanced. AI attained accuracy in objective information and promptly rectified the error.