PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2026Turkiye Klinikleri Journal of Dental Sciences0 citationsOpen Access

Comparison of Performance of Leading Large Language Models in Answering Medical Pathology Questions in Dentistry Specialization Education Entrance Exams: A Cross-Sectional Research

ÖEÖmer EkiciİÇİsmail ÇALIŞKAN

Key Points

  • This study aims to evaluate and compare the accuracy of AI large language models in answering pathology questions in dentistry specialization education exams.
  • Analyzed 52 pathology questions from 13 exams published by the official Student Selection and Placement Center.
  • Questions posed to different LLMs simultaneously by a single operator.
  • Used chi-square analysis to compare correct response rates across LLMs.
  • Correct answer rates were highest for ChatGPT-4o (100%) and lowest for Co-Pilot (76.92%).
  • Basic pathology questions had higher accuracy rates than clinical pathology questions.
  • A significant difference in accuracy was observed between clinical pathology and overall questions (p<0.05).

Abstract

Objective: Artificial intelligence (AI) based large language models (LLMs) have recently become an effective and efficient tool in education and learning. The purpose of this study is to comparatively evaluate the accuracy of the answers given by different AI-supported chatbots to the Medical Pat hology questions asked in the Dentistry Specialization Education Entrance Exams (DSE). Material and Methods: A total of 52 pathology questions from 13 exams published on the official website of Student Selection and Placement Center (Öğrenci Seçme ve Yerleştirme Merkezi) were included in the study. Questions were directed simultaneously to LLMs by a single operator. Chi-square analysis was used to compare correct response rates among LLMs in all questions. Results: The order of correct answer rates of LLMs to all questions was as follows: ChatGPT-4o (100%), Chat GPT4 (96.15%), Gemini 2.0 (90.38%) and Claude 3 Sonnet (90.38%), Gemini 1.5 (86.53%), Co-pilot (76.92%). In general, correct answer percentages of LLMs in basic pathology questions were higher than in clinical pathology questions. While no statistically significant difference was observed between correct answers of LLMs to basic pathology questions (p=0.542), a significant difference was observed between correct answers of LLMs in clinical pathology and all questions (p<0.05). Conclusion: In this study, the highest accuracy rate was found in GPT-4o and the lowest rate was found in Co-Pilot. The findings show that LLMs have the potential to be used as a supportive tool for students and academics in pathology education.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ekici et al. (2026) studied this question.

synapsesocial.com/papers/6a095b1b7880e6d24efe0d65https://doi.org/10.5336/dentalsci.2025-111376
Ask AI
Helpful
Bookmark
Share
View Full Paper