Comparative diagnostic agreement of a supervised machine learning model and a general-purpose, zero-shot, non-domain-adapted large language model for classifying headache disorders using structured questionnaires.

Masahito Katsuki, Kieran Moran, Siobhán O'Connor, Tomás Ward, Yutaro Fuse, Omid Kohandel Gargari, Marina Romozzi, Alicia Gonzalez-Martinez, Miguel Á Huerta, Woo-Seok Ha, Jackson Ts Cheung, Yasuhiko Matsumori

Journal: Cephalalgia : an international journal of headache 2026;46(4):3331024261441574

PMID: 42635070

Abstract

BackgroundAccurate diagnosis of headache disorders is essential in clinical practice. Supervised machine learning models trained on structured clinical data have shown good performance, whereas the diagnostic ability of large language models (LLMs) for headache disorders has not been evaluated. This study compared a validated machine learning classifier with a general-purpose, zero-shot, non-domain-adapted LLM using the same structured patient questionnaire data, focusing on their agreement with specialist-confirmed diagnoses as the ground truth. This study was designed to reflect current real-world use scenarios, in which clinicians may apply off-the-shelf LLMs for diagnostic purposes without few-shot prompting, domain-specific fine-tuning, or adaptation, rather than to assess the theoretical upper limits of LLM capabilities.MethodsWe analyzed 1818 patients from an independent hold-out test cohort who completed a 22-item structured headache questionnaire and received specialist-confirmed diagnoses. A previously developed machine learning model and a general-purpose, non-domain-adapted LLM (GPT-4.1 with zero-shot prompting) each generated five-class International Classification of Headache Disorders, 3rd edition (ICHD-3)-based predictions: migraine and/or medication-overuse headache (MOH), tension-type headache (TTH), trigeminal autonomic cephalalgias (TACs), other primary headache disorders, and secondary headaches. Agreement with the specialist's diagnosis and diagnostic performance metrics were calculated. Class-wise sensitivity and specificity were compared using McNemar's test.ResultsThe machine learning classifier showed significantly higher diagnostic agreement with the specialist than the LLM (Cohen's κ: 0.46 vs. 0.26; 95% confidence interval of the difference: 0.15-0.25). Although the LLM showed slightly higher macro-averaged sensitivity (balanced accuracy) than the machine learning model, the machine learning classifier showed higher macro-averaged precision, specificity, and F-value. Class-wise analysis showed that the machine learning model demonstrated greater sensitivity for migraine and/or MOH and secondary headaches, while the LLM showed higher sensitivity for TTH. Regarding specificity, the machine learning model outperformed the LLM in TTH, TACs, and other primary headache disorders, whereas the LLM showed higher specificity only for migraine and/or MOH.ConclusionsA supervised machine learning model trained on real-world clinical data showed better agreement with a specialist-confirmed diagnosis than a general-purpose, zero-shot, non-domain-adapted LLM. These findings indicate that, in its current off-the-shelf configuration under this experimental setting, the diagnostic agreement between a general-purpose LLM and specialists can be limited for headache disorders.

Address: Physical Education and Health Center, Nagaoka University of Technology, Nagaoka, Japan.; School of Health and Human Performance, Dublin City University, Dublin, Ireland.; Insight Research Ireland Centre for Data Analytics, Dublin City University, Dublin, Ireland.; Department of Biostatistics, Graduate School of Medicine, Saitama Medical University, Saitama, Japan.; School of Health and Human Performance, Dublin City University, Dublin, Ireland.; Insight Research Ireland Centre for Data Analytics, Maynooth University, Kildare, Ireland.; Department of Sport Science and Nutrition, Maynooth University, Kildare, Ireland.; School of Health and Human Performance, Dublin City University, Dublin, Ireland.; Insight Research Ireland Centre for Data Analytics, Dublin City University, Dublin, Ireland.; School of Computing, Dublin City University, Dublin, Ireland.; Department of Artificial Intelligence Medicine, Graduate School of Medicine, Chiba University, Chiba, Japan.; Headache Department, Iranian Center of Neurological Research, Neuroscience Institute, Tehran University of Medical Sciences, Tehran, Iran.; Dipartimento Universitario di Neuroscienze, Università Cattolica del Sacro Cuore, Rome, Italy.; Neurologia, Dipartimento di Scienze dell'invecchiamento, Neurologiche, Ortopediche e della Testa-Collo, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Rome, Italy.; Neurology Department, La Princesa University Hospital, Madrid, Spain.; Translational Research Group in Multimodal Biomarkers in Neurological Diseases, IIS-Princesa, Madrid, Spain.; Autonomous University, Madrid, Spain.; Department of Pharmacology, University of Granada, Granada, Spain.; Biosanitary Research Institute ibs.GRANADA, Granada, Spain.; Department of 17harmacology, University of Cambridge, Cambridge, UK.; Department of Neurology, Severance Hospital, Yonsei University College of Medicine, Seoul, Korea.; UCL Faculty of Medical Sciences, London, UK.; Sendai Headache and Neurology Clinic, Sendai, Japan.
Bant logo

© Copyright 2026, Nutrition Evidence

NED wishes to thank the following organisations for their support:

We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.