Agreement of Feline Grimace Scale scores between chatbots and an expert rater
摘要
The performance of large language models (LLMs) during acute pain assessment in cats has not been evaluated. This study evaluated the agreement of the Feline Grimace Scale (FGS) scoring between four chatbots (ChatGPT, Gemini, Claude AI, and Perplexity) and an expert veterinarian, including bias and limits of agreement (LoA), and whether bias would be reduced when retested after two months. Fifty cat facial images were scored twice, two months apart, by each chatbot using the FGS (ear position, orbital tightening, muzzle tension, whiskers change and head position). The Bland–Altman method was used to analyze bias and limits of LoA. Chatbots showed positive bias, indicating underestimation of FGS scores. Claude AI presented an acceptable bias (< 0.1) suggesting good agreement after retesting. However, its LoA spanned the FGS threshold for analgesia (0.39). The LoA of ChatGPT did not span the threshold, but presented unacceptable bias (> 0.1). Gemini showed unacceptable bias and LoA spanned the FGS threshold. Perplexity showed unacceptable bias and its LoA spanned the threshold after retesting. Most chatbots showed poor agreement and could have compromised analgesia during testing and/or retesting; pain scoring could be overestimated or underestimated to an extent that would cause overtreatment or undertreatment of pain.