Background/aim <p>This study aims to explore ChatGPT’s ability to provide comprehensive information on AIH, its potential role in supporting the diagnostic and therapeutic processes for the disease and its broader implications for the use of artificial intelligence in healthcare.</p> Materials and methods <p>A total of 45 questions were designed and organized into 3 groups, each consisting of 15 questions (Group 1, Group 2, Group 3). The questions were categorized based on diffuculty, with Group 1 being the easiest and Group 3 the most challenging. Additionally, 5 case-based questions were formulated and ChatGPT was asked to provide diagnoses for these cases. All questions were re-asked after 14 days to evaluate the short-term stability and consistency of the artificial intelligence responses.</p> Results <p>There was no statistically significant difference in the responses provided by ChatGPT initially (<i>p</i> = 0.328). However, when the questions were re-asked after 14 days, a significant difference was observed between the groups in terms of accuracy and completeness (<i>p</i> = 0.045 and <i>p</i> = 0.015, respectively). This difference was primarily due to the lower accuracy level in the responses for group 3 questions. Furthermore, when comparing ChatGPT’s responses at 2-week intervals, a statistically significant improvement was noted, except for the completeness score of group 3 questions and the case-based questions.</p> Conclusion <p>Our findings suggest that ChatGPT may play a greater role in patient education and in supporting physicians in disease management. Future studies that test AI-generated responses from multiple perspectives and involve more centers could provide clearer guidance on this topic.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of the effectiveness of ChatGPT in supporting the management of autoimmune hepatitis effectiveness of ChatGPT in autoimmune hepatitis

  • Muhammed Mustafa İnce,
  • Ersan Özaslan

摘要

Background/aim

This study aims to explore ChatGPT’s ability to provide comprehensive information on AIH, its potential role in supporting the diagnostic and therapeutic processes for the disease and its broader implications for the use of artificial intelligence in healthcare.

Materials and methods

A total of 45 questions were designed and organized into 3 groups, each consisting of 15 questions (Group 1, Group 2, Group 3). The questions were categorized based on diffuculty, with Group 1 being the easiest and Group 3 the most challenging. Additionally, 5 case-based questions were formulated and ChatGPT was asked to provide diagnoses for these cases. All questions were re-asked after 14 days to evaluate the short-term stability and consistency of the artificial intelligence responses.

Results

There was no statistically significant difference in the responses provided by ChatGPT initially (p = 0.328). However, when the questions were re-asked after 14 days, a significant difference was observed between the groups in terms of accuracy and completeness (p = 0.045 and p = 0.015, respectively). This difference was primarily due to the lower accuracy level in the responses for group 3 questions. Furthermore, when comparing ChatGPT’s responses at 2-week intervals, a statistically significant improvement was noted, except for the completeness score of group 3 questions and the case-based questions.

Conclusion

Our findings suggest that ChatGPT may play a greater role in patient education and in supporting physicians in disease management. Future studies that test AI-generated responses from multiple perspectives and involve more centers could provide clearer guidance on this topic.