PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 14, 2026Journal of Oral Pathology and Medicine0 citations

Evaluation of GPT ‐5, a Large Language Model, in Replicating German Clinical Practice Guideline Recommendations in Oral Oncology: A Cross‐Sectional Concordance Study

View Full Paper
JHJulius HirschKSKeskanya SubbalekhaCTChatpong Tangmanee

Key Points

  • This study aims to evaluate how accurately the language model GPT-5 reproduces recommendations from German clinical practice guidelines for oral oncology.
  • Conducted a cross-sectional analysis comparing GPT-5 outputs to German clinical practice guidelines.
  • Guideline statements were input into GPT-5, which confirmed or rejected the statements.
  • Inverted versions of the recommendations were also tested for methodological robustness.
  • Concordance and accuracy were measured using Cohen's 𝝹 statistic.
  • GPT-5 perfectly affirmed all true recommendations from the guidelines.
  • When both original and inverted statements were assessed together, the agreement was very high (𝝹 = 0.96).
  • Complete concordance was noted between GPT-5 responses and the original guideline statements (𝝹 = 1.0).
  • The majority of references in the guidelines were in English and from outside Germany.

Abstract

ABSTRACT Background Artificial intelligence (AI) technologies, particularly large language models (LLMs) such as ChatGPT, are increasingly utilised in medical education and clinical information retrieval. Nevertheless, their capacity to accurately reproduce recommendations from established clinical practice guidelines (CPGs) has not been thoroughly examined. The present study evaluated the concordance between responses generated by GPT‐5 and recommendations contained in German CPGs addressing oral potentially malignant disorders (OPMDs) and oral carcinomas (OCs). Methods A cross‐sectional analytical comparison was performed between GPT‐5 outputs and German CPG recommendations available as of October 2025. Individual guideline statements were entered verbatim into GPT‐5, which was asked to confirm or reject the statements. To assess methodological robustness, inverted versions of the same statements were additionally tested. GPT‐5 was accessed through the free version without internet connectivity to ensure that responses originated solely from the model's internal training data. Accuracy was defined as the proportion of correctly classified statements. Concordance between guideline content and model responses was quantified using Cohen's 𝝹. Results Two German CPGs comprising 111 recommendations were included: the S2k guideline for OPMDs (15 recommendations) and the S3 guideline for OCs (96 recommendations). GPT‐5 correctly affirmed all authentic recommendations and rejected all inverted statements. Agreement between guideline statements and GPT‐5 responses was perfect when the original recommendations were analysed (𝝹 = 1.0) and remained very high when both original and inverted statements were evaluated jointly (𝝹 = 0.96). The majority of references cited within the guidelines were published in English (> 93%) and originated from outside Germany (> 77%). Conclusion When guideline recommendations were presented verbatim, GPT‐5 demonstrated complete concordance with German oral oncology CPGs. These findings indicate that the model is capable of recognising and retrieving established guideline information. However, this experimental design evaluates recognition of existing statements rather than autonomous clinical reasoning. At present, LLMs should therefore be regarded primarily as educational and informational tools rather than a replacement for expert clinical judgement in oral oncology.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hirsch et al. (2026) studied this question.

synapsesocial.com/papers/69ddd975e195c95cdefd6d93https://doi.org/10.1111/jop.70140
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Feasibility and Concordance of a Large Language Model (ChatGPT-5) as a Clinical Decision Support Tool in Gynecologic Oncology Tumor Boards: A Blinded, Multi-Observer Study2026
  2. 2Feasibility and concordance of a large language model (ChatGPT-5) as a clinical decision support tool in gynecologic oncology tumor boards: A blinded, multi-observer study.2026
  3. 3Assessing the performance of ChatGPT in addressing ethical dilemmas in oncology.2026
  4. 4GPT-4 for Information Retrieval and Comparison of Medical Oncology Guidelines2024 · 87 citations
  5. 5Performance of leading large language models in adhering to clinical guidelines for anaplastic thyroid cancer: a comparative study2026