PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 2026Journal of International Medical Research1 citationsOpen Access

Accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society left atrial appendage occlusion guidelines

View Full Paper
YJYunpeng JinCFChao FengJZJianqiang Zhao

Key Points

  • The research aimed to evaluate the accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 cardiology guidelines.
  • Responses from four large language models were evaluated for eight clinical questions.
  • Cardiologists rated accuracy using a six-point Likert scale.
  • Reproducibility and inter-rater reliability were assessed with Fleiss' kappa and intraclass correlation coefficients.
  • ChatGPT-5 had a median accuracy of 5.5 compared to 5 for DeepSeek models.
  • ChatGPT models showed significantly higher accuracy than DeepSeek models (p < 0.001).
  • Reproducibility was excellent for ChatGPT-5 (0.803), while DeepSeek showed lower values.

Abstract

ObjectiveTo evaluate the accuracy of ChatGPT and DeepSeek in answering guideline-based clinical questions in cardiology.MethodsIn August 2025, responses generated from four large language models to eight clinical questions based on the 2025 Society for Cardiovascular Angiography (b) more incorrect than correct; (c) nearly equally correct and incorrect; (d) more correct than incorrect; (e) nearly all correct; and (f) completely correct. Reproducibility (Fleiss' kappa coefficient, five repeated queries) and inter-rater reliability (intraclass correlation coefficient) were assessed.ResultsThe median (interquartile range) accuracy scores were 5.5 (5, 6) for ChatGPT-5, 6 (5, 6) for ChatGPT-4o, and 5 (4, 6) for both DeepSeek-R1 and DeepSeek-V3, with a significant overall difference (p < 0.001). Pairwise comparisons showed significantly higher accuracy for ChatGPT models than for DeepSeek models (all p < 0.001), whereas no significant differences were observed between ChatGPT-5 and ChatGPT-4o (p = 0.518) or between DeepSeek-R1 and DeepSeek-V3 (p = 0.812). Reproducibility (Fleiss' kappa coefficient) was excellent for ChatGPT-5 (0.803) and good for ChatGPT-4o (0.574), DeepSeek-R1 (0.577), and DeepSeek-V3 (0.618). Overall inter-rater reliability was moderate (intraclass correlation coefficient = 0.463).ConclusionsChatGPT and DeepSeek demonstrated high accuracy and reproducibility but moderate inter-rater reliability, necessitating further validation for educational use.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jin et al. (2026) studied this question.

synapsesocial.com/papers/69e07c972f7e8953b7cbdd17https://doi.org/10.1177/03000605261438348
Ask AI
Helpful
Bookmark
Share
View Full Paper