From GPT-4 to Expert-Endorsed Athlete Guidance: A Delphi Consensus on Sleep and Jet Lag
摘要
Large language models (LLMs) including GPT-4 are increasingly used to generate health information, but concerns persist about their accuracy and relevance, particularly for elite athletes.
ObjectiveThis study used GPT-4-generated frequently asked question (FAQ) responses on sleep and jet lag as the starting material for expert evaluation and consensus development, with the goal of producing consensus-based, athlete-specific guidance while identifying the limitations of AI-generated content.
MethodsBetween November 2024 and March 2025, n = 17 international sleep and circadian experts from the Athlete Travel & Sleep Interest Group (ATSIG) participated in a two-round Delphi process. Experts rated 20 GPT-4-generated FAQ responses (10 on sleep, 10 on jet lag) for appropriateness using a 6-point Likert scale and provided qualitative feedback. Items were revised after round 1 using inductive thematic coding. Consensus was defined as ≥ 70% of participants rating an item as appropriate (scores 5–6) and ≤ 15% as non-appropriate (scores 1–2). Statistical analyses included Wilcoxon signed-rank tests, convergence metrics and dissent detection (outlier and bipolarity analysis).
ResultsIn round 1, 15 of 20 items (75%) met consensus; by round 2, 18 of 20 (90%) achieved consensus. For sleep items, 7 of 10 reached consensus in round 1 and 9 in round 2; for jet lag, 8 items reached consensus in round 1 and 9 in round 2. Sleep Q6 (sleep and injury risk) narrowly missed the consensus threshold with 64.7% agreement, while Jet Lag Q9 (melatonin and sleep aids) remained below the 70% threshold. No item showed bimodal score distributions, suggesting no polarized disagreement. Descriptive rating patterns, increased consensus and qualitative expert feedback indicated improved clarity, accuracy and athlete-specific relevance after the Delphi process, although item-level statistical comparisons did not remain significant after Bonferroni correction for multiple testing. Qualitative analysis identified common concerns: for sleep items—imprecise or misleading content (55%), lack of athlete-specific relevance (30%) and outdated evidence (11%); for jet lag—outdated evidence (36%), imprecise or misleading content (34%) and formatting issues (17%).
ConclusionsGPT-4-derived content may serve as useful preliminary material for expert discussion but should not be used as standalone guidance. Expert evaluation improved the clarity, safety and athlete-specific relevance of most sleep and jet-lag responses, while the final outputs should be interpreted as consensus-based guidance rather than definitive proof of correctness.