ABSTRACT Background Crohn's disease (CD) requires accurate, guideline‐based diagnosis and management, yet the growing complexity of therapeutic options and international recommendations creates challenges for timely clinical decision‐making. Large language models (LLMs) have emerged as potential tools for clinical decision support, but their performance in guideline‐based gastroenterology remains insufficiently characterized. Methods We conducted a structured comparative evaluation of ChatGPT 5.0 and DeepSeek R1 using 47 standardized clinical questions derived from the 2023 Chinese guidelines for the diagnosis and treatment of CD. Each question was submitted independently to both models under identical conditions, without iterative clarification. Responses were assessed across multiple dimensions, including coverage, guideline consistency, agreement, response length, detail coverage, readability, and structural clarity. Domain‐specific analyses were performed for diagnostic, therapeutic, and surgical/management questions. Results Both models demonstrated high alignment with guideline‐based recommendations, with guideline consistency rates of 93% for ChatGPT 5.0% and 96% for DeepSeek R1, and an inter‐model agreement rate of approximately 87%. ChatGPT 5.0 showed complete detail coverage across diagnostic, therapeutic, and surgical/management domains (100%) and generated shorter, more concise responses (mean 208 words; range 142–300) with higher readability (mean 4.6/5). DeepSeek R1 achieved 87.5% overall detail coverage, with slightly lower performance in diagnostic (88.2%) and therapeutic (85.2%) categories but equivalent performance in surgical/management questions. Its responses were longer (mean 291 words; range 128–417) and incorporated more trial data, drug comparisons, and cross‐guideline references, although readability was lower (mean 3.4/5). Conclusions ChatGPT 5.0 and DeepSeek R1 demonstrated distinct yet complementary strengths in guideline‐based CD applications. ChatGPT 5.0 functioned more effectively as a clinical‐ready assistant by providing concise, highly readable, and actionable outputs, whereas DeepSeek R1 was better suited to academic and research‐oriented contexts through deeper evidence integration. A hybrid workflow combining rapid point‐of‐care support from ChatGPT 5.0 with evidence synthesis from DeepSeek R1 may offer optimal utility. Further real‐world validation is needed before routine clinical implementation.
Zhou et al. (2026) studied this question.