Assessing Large Language Models for Patient Interaction in Orthodontics: A Brief Evaluation on Obstructive Sleep Apnea Queries
摘要
Obstructive Sleep Apnea (OSA) remains underdiagnosed despite its high prevalence and health risks. With increasing patient reliance on artificial intelligence (AI) for health guidance, this study aimed to evaluate the accuracy, comprehensiveness, and empathy of responses generated by large language models (LLMs) to OSA-related queries, as judged by orthodontists.
MethodsTen commonly asked OSA-related patient questions were posed to five LLMs (ChatGPT-3.4, ChatGPT-4.0, LLaMA, Bing, and Gemini). Responses were collected and anonymized, then evaluated by 20 orthodontists with over ten years of clinical experience. Ratings were assigned using a standardized 4-point scale (accuracy, comprehensiveness, empathy, and clinical relevance). Statistical analysis was conducted using the Kruskal–Wallis test with Bonferroni post hoc correction.
ResultsSignificant differences were observed across models (p < 0.05). LLaMA consistently outperformed other models, offering comprehensive, empathetic, and clinically relevant responses. Bing was valued for citation-based outputs, while Gemini was appreciated for its disclaimers. ChatGPT models performed less favourably across multiple domains.
ConclusionLLaMA demonstrated superior response quality in OSA-related interactions; however, AI tools should be considered complementary resources rather than substitutes for professional diagnosis and care. Integration with clinician oversight remains essential for patient safety.