<p>With growing reliance on AI chatbots for parenting support, this study presents the first evaluation of large language models (LLMs) in addressing common autism-related questions. It compared ChatGPT, Google Gemini, and DeepSeek based on the accuracy, clarity, and usefulness of their responses. The findings aim to inform parents and clinicians about the strengths and limitations of using AI tools in early ASD care. Twenty common questions about Autism Spectrum Disorder (ASD) were identified through content analysis of social media, Google Trends, and ASD forums. These questions were refined by two educational psychologists, and standardized benchmark answers were created by a panel of pediatric neurodevelopment specialists. Two blinded pediatric autism experts then evaluated the AI-generated responses based on quality, as well as usefulness and reliability. GPT-4 achieved the highest mean quality score (M = 4.85, SD = 0.36), followed by Gemini and DeepSeek (both M = 4.55, SD = 0.51; <i>p</i> &gt; 0.05). For usefulness, GPT-4 scored M = 6.40 (SD = 0.75), Gemini M = 6.10 (SD = 0.85), and DeepSeek M = 6.05 (SD = 0.82; <i>p</i> &gt; 0.05). In reliability ratings, Gemini led with M = 6.40 (SD = 0.82), GPT-4&#xa0;M = 6.25 (SD = 0.71), and DeepSeek M = 5.95 (SD = 0.94; <i>p</i> &gt; 0.05). Findings indicated that AI-based chatbots, by providing rapid, comprehensible, and evidence-based guidance on early signs, interventions, and family support, demonstrate significant potential in bridging the information gap for parents—especially when access to specialists is limited.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing AI-Based Large Language Models (ChatGPT, Google Gemini, and DeepSeek) for Common Parent Questions about Autism: Acceptability, Readability, and Accuracy

  • Abdullah Ahmed Almulla,
  • Mohamad Ahmad Saleem Khasawneh

摘要

With growing reliance on AI chatbots for parenting support, this study presents the first evaluation of large language models (LLMs) in addressing common autism-related questions. It compared ChatGPT, Google Gemini, and DeepSeek based on the accuracy, clarity, and usefulness of their responses. The findings aim to inform parents and clinicians about the strengths and limitations of using AI tools in early ASD care. Twenty common questions about Autism Spectrum Disorder (ASD) were identified through content analysis of social media, Google Trends, and ASD forums. These questions were refined by two educational psychologists, and standardized benchmark answers were created by a panel of pediatric neurodevelopment specialists. Two blinded pediatric autism experts then evaluated the AI-generated responses based on quality, as well as usefulness and reliability. GPT-4 achieved the highest mean quality score (M = 4.85, SD = 0.36), followed by Gemini and DeepSeek (both M = 4.55, SD = 0.51; p > 0.05). For usefulness, GPT-4 scored M = 6.40 (SD = 0.75), Gemini M = 6.10 (SD = 0.85), and DeepSeek M = 6.05 (SD = 0.82; p > 0.05). In reliability ratings, Gemini led with M = 6.40 (SD = 0.82), GPT-4 M = 6.25 (SD = 0.71), and DeepSeek M = 5.95 (SD = 0.94; p > 0.05). Findings indicated that AI-based chatbots, by providing rapid, comprehensible, and evidence-based guidance on early signs, interventions, and family support, demonstrate significant potential in bridging the information gap for parents—especially when access to specialists is limited.