PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 10, 2026SLEEP0 citations

0357 Machine Learning Meets Medicine: Assessment of Large Learning Modules Answers to Patient Sleep-Related Questions

View Full Paper
FDFauzieh DabajaMSMehwish SajidXTXinhang Tu

Key Points

  • This study evaluates the accuracy and clinical appropriateness of answers from three large learning modules regarding sleep-related medical questions.
  • Prompt generated for common sleep medicine questions asked by users.
  • Responses from ChatGPT 5.1, Gemini 3, and Perplexity were independently evaluated by five board-certified sleep medicine physicians.
  • Friedman test was applied to compare accuracy and clinical appropriateness scores across the three models.
  • Average accuracy scores: ChatGPT 80%, Gemini 80%, Perplexity 87%.
  • Clinical appropriateness scores averaged 85% (ChatGPT), 79% (Gemini), 88% (Perplexity).
  • Statistically significant difference found in clinical appropriateness ratings (χ2 = 6.87, p = 0.0323).

Abstract

Abstract Introduction Large learning modules (LLM) are now available for use by the general population, and patients now approach these modules to answer medical questions, despite or in lieu of vising a physician. These modules provide accessible, easy to understand, and rapid results. The accuracy of these modules varies, and their accuracy and clinical alignment remain uncertain. This study aimed to evaluate and compare three different LLM answers about sleep medical problems. Methods A prompt was generated to request a set of most common sleep medicine asked questions by ChatGPT users. Questions then entered into three Large Learning Modules: ChatGPT 5.1, Gemini 3 and Perplexity. Answers were reviewed independently by five board certified sleep medicine physicians for accuracy and clinical appropriateness based on a 5-point scale, reviewers were blinded for the names of the LLM used. Friedmans test was used to compare rating across three platforms treating each answer as a repeated measure. Results Across the 10 standardized sleep-related questions, the average accuracy scores for the three large language models (ChatGPT, Gemini, and Perplexity) were 80%, 80%, and 87%, respectively. Clinical appropriateness scores averaged 85% for ChatGPT, 79% for Gemini, and 88% for Perplexity. A Friedman test comparing accuracy ratings across models demonstrated no statistically significant difference (χ2 = 5.20, p = 0.0743). In contrast, the Friedman test for clinical appropriateness revealed a statistically significant difference among the three platforms (χ2 = 6.87, p = 0.0323), indicating meaningful variation in the clinical suitability of responses. Conclusion Clinicians should be aware that patients frequently use learning modules to help answer clinical concerns. Although they may help enhance patient engagement they may be laced with subtle inaccuracies, over generalized recommendations, and false reassurance. Clinical encounters may involve clarifying misconceptions and helping guide patients to more individualized and evidence-based care. LLM can vary in their clinical appropriateness and as more models are being trained to use more accurate medical resources this will improve its output. Support (if any)

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dabaja et al. (2026) studied this question.

synapsesocial.com/papers/6a00217ac8f74e3340f9c4f5https://doi.org/10.1093/sleep/zsag091.0357
Ask AI
Helpful
Bookmark
Share
View Full Paper