
Lead
Generative AI tools like ChatGPT are increasingly used by athletes, coaches, and support staff as a quick source of health and performance advice. But how reliable is that advice when it comes to the complex, evidence-based domain of sleep science and jet lag management in elite sport? A new study published in Sports Medicine put GPT-4 to the test, using a structured Delphi consensus process to have international sleep experts evaluate and refine AI-generated answers on 20 key topics. The result is a mixed verdict: GPT-4 offers a useful starting point, but it is not yet ready to stand alone as authoritative guidance for athletes.
The study, led by Jacopo Vitale and a team that includes Alan McCall, Shona Halson, and Dina C Janse van Rensburg, was conducted under the auspices of the Athlete Travel & Sleep Interest Group (ATSIG). It is one of the first systematic attempts to benchmark large language model output against expert consensus in the sports medicine literature.
What they found
The research team generated a set of FAQ-style responses on sleep and jet lag using GPT-4, covering 20 items that reflect common questions athletes and practitioners ask. These items ranged from the basic physiology of sleep and circadian rhythms to practical strategies for managing jet lag, napping, sleep hygiene, and recovery. The AI-generated content was then presented to 17 international experts in sleep and circadian science, who evaluated each item for accuracy, completeness, and suitability for an athletic population.
The evaluation followed a two-round Delphi method, a well-established technique for building consensus among experts through iterative, anonymous feedback. After Round 1, 15 out of 20 items (75%) had reached a predefined consensus threshold. After incorporating expert feedback and revising the content, Round 2 pushed that number to 18 out of 20 items (90%), indicating strong overall agreement among the panel.
Two items failed to reach consensus in either round. The first was the relationship between sleep and injury risk, which received only 64.7% agreement. The second was guidance on melatonin and other sleep aids. These topics are inherently contentious in sports medicine, where the evidence base is still evolving and expert opinions diverge on appropriate recommendations.
Beyond the consensus data, the study catalogued expert concerns about GPT-4’s original answers. These concerns fell into several categories. The most prominent was that the content was imprecise or misleading: 55% of experts raised this concern for sleep-related items, and 34% did so for jet lag items. A lack of athlete-specific contextualization was flagged by 30% of the panel. Additionally, experts noted that some of the content relied on outdated evidence, with rates ranging from 11% to 36% depending on the topic.
Why it matters
Elite athletes operate in an environment where small margins separate success from failure. Sleep quality and circadian alignment are recognized as critical factors in recovery, cognitive function, injury prevention, and competitive performance. Yet access to reliable, athlete-specific sleep guidance remains uneven, particularly for traveling athletes who face frequent time zone shifts. Generative AI tools promise to democratize access to expert-level knowledge, but their outputs are only as good as the data they were trained on and the validation they receive.
This study matters because it provides a concrete, reproducible method for bridging the gap between AI-generated information and expert-endorsed guidance. Rather than treating GPT-4 output as a finished product, the researchers used it as preliminary material that could be refined through structured expert review. The Delphi process identified specific weaknesses in the AI content and systematically improved them. The result is a set of 18 expert-endorsed FAQ responses that carry significantly more authority than raw GPT-4 output alone.
The findings also have implications beyond sports medicine. The same pattern likely holds in other specialized domains where precision, currency, and contextual nuance matter: clinical guidelines, nutrition advice, training programming, and rehabilitation protocols. AI can accelerate the drafting process, but human expertise remains essential for vetting and refining the output.
For practitioners, the study reinforces a practical caution: do not take AI-generated health and performance advice at face value. The 55% rate of expert concern about imprecision in sleep content, combined with the failure to reach consensus on melatonin and injury risk, shows that GPT-4 can produce answers that sound authoritative but are not sufficiently accurate or complete for real-world application with athletes.
Limits
The study has several limitations worth noting. The sample of 17 experts, while diverse and internationally recognized, is relatively small. A larger panel might have yielded different consensus thresholds or surfaced additional concerns. The Delphi method itself, while rigorous, is time-intensive and depends on the quality of the initial materials and the engagement of the participants. The GPT-4 output used in the study reflects a specific snapshot of the model’s capabilities; given that large language models are updated frequently, the results may not generalize to newer versions or to other AI systems.
The scope of the FAQ items was also limited to 20 topics selected by the research team. There are many additional questions athletes might ask about sleep and jet lag that were not included, and the study does not assess how well the expert-refined answers perform in real-world use or whether they improve athlete outcomes. Finally, the study did not compare GPT-4 output against that of other AI models, so it is unclear whether alternative systems would fare better or worse.
Bottom line
GPT-4-generated content on sleep and jet lag for athletes can serve as useful preliminary material but requires expert refinement before it is suitable for guidance. The Delphi consensus process improved agreement from 75% to 90%, and the resulting 18 expert-endorsed items represent a meaningful step toward trustworthy AI-assisted guidance in sports medicine. However, two important topics (sleep and injury risk, melatonin and sleep aids) could not achieve consensus even after two rounds, underscoring that some areas remain too contentious or evidence-poor for any single source to resolve. Athletes, coaches, and practitioners should treat AI-generated sleep advice as a draft, not a final answer, and seek expert verification for high-stakes decisions.
Source
Vitale J, McCall A, Halson S, Janse van Rensburg DC, et al. From GPT-4 to Expert-Endorsed Athlete Guidance: A Delphi Consensus on Sleep and Jet Lag. Sports Medicine. 2026. DOI: 10.1007/s40279-026-02484-7. PMID: 42470603.

