Validation of an AI-Based platform for structured diagnosis of headache disorders using ICHD-3 criteria
摘要
The diagnosis of headache disorders remains a clinical challenge, particularly for non-specialists, due to the complexity of the International Classification of Headache Disorders, 3rd edition (ICHD-3), and the absence of biomarkers. Large language models (LLMs) represent a promising tool to support accurate and scalable diagnostic classification, especially in resource-limited settings.
ObjectiveTo validate the performance of a free, multilingual clinical decision support platform—Head.AI—designed to classify headache cases using GPT-4o and a structured implementation of ICHD-3.
MethodsWe conducted an independent validation using 315 expert-generated vignettes representing 215 ICHD-3 diagnoses, input into Head.AI and three other platforms (Claude Sonnet 4.0, Grok 3.0, and Gemini 2.5). Outcomes included diagnostic accuracy (rank of correct diagnosis), calibration, and citation rate.
ResultsThe algorithm correctly identified the top diagnosis in 89.5% of cases (vs. 74–80% in comparators), with a citation rate >97% and calibration (Brier score 0.153). It maintained consistent performance across primary and secondary headaches and achieved first-hypothesis accuracy >74% in difficult cases. Logistic regression confirmed Head.AI had significantly higher odds of correct classification (ORs vs. comparators: 2.04–2.86; all p < 0.01).
ConclusionOur algorithm demonstrated high diagnostic accuracy across a broad spectrum of headache disorders, exceeding the performance reported in prior studies, though direct comparison should be interpreted with caution due to methodological differences. Its public availability, structured knowledge base, and educational potential make it a valuable contribution to AI-assisted headache care. The platform is freely accessible at www.head-ai.com.br.