Abstract Introduction Large learning modules (LLM) are now available for use by the general population, and patients now approach these modules to answer medical questions, despite or in lieu of vising a physician. These modules provide accessible, easy to understand, and rapid results. The accuracy of these modules varies, and their accuracy and clinical alignment remain uncertain. This study aimed to evaluate and compare three different LLM answers about sleep medical problems. Methods A prompt was generated to request a set of most common sleep medicine asked questions by ChatGPT users. Questions then entered into three Large Learning Modules: ChatGPT 5.1, Gemini 3 and Perplexity. Answers were reviewed independently by five board certified sleep medicine physicians for accuracy and clinical appropriateness based on a 5-point scale, reviewers were blinded for the names of the LLM used. Friedmans test was used to compare rating across three platforms treating each answer as a repeated measure. Results Across the 10 standardized sleep-related questions, the average accuracy scores for the three large language models (ChatGPT, Gemini, and Perplexity) were 80%, 80%, and 87%, respectively. Clinical appropriateness scores averaged 85% for ChatGPT, 79% for Gemini, and 88% for Perplexity. A Friedman test comparing accuracy ratings across models demonstrated no statistically significant difference (χ2 = 5.20, p = 0.0743). In contrast, the Friedman test for clinical appropriateness revealed a statistically significant difference among the three platforms (χ2 = 6.87, p = 0.0323), indicating meaningful variation in the clinical suitability of responses. Conclusion Clinicians should be aware that patients frequently use learning modules to help answer clinical concerns. Although they may help enhance patient engagement they may be laced with subtle inaccuracies, over generalized recommendations, and false reassurance. Clinical encounters may involve clarifying misconceptions and helping guide patients to more individualized and evidence-based care. LLM can vary in their clinical appropriateness and as more models are being trained to use more accurate medical resources this will improve its output. Support (if any)
Dabaja et al. (2026) studied this question.