PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 20, 2026BMC Oral Health0 citationsOpen Access

Comparison of the diagnostic accuracy of dentists and ChatGPT in jawbone lesions

NSNezahat Sena SatanUPUmut PamukçuBTBarış Erkut Türk

Key Points

  • This study evaluates the diagnostic accuracy of ChatGPT for jawbone lesions and compares it with dentists and specialists.
  • Selected thirty cases with clinical information, panoramic radiographs, and histopathological diagnoses.
  • Collected data through a questionnaire distributed electronically to OMFR, OMFS, and general dentists.
  • Analyzed data using Wilcoxon Signed Rank, Mann–Whitney U, and Kruskal–Wallis tests.
  • ChatGPT had a diagnostic accuracy of 46.67%.
  • OMFR and OMFS achieved significantly higher accuracy rates (67.71% and 58.96%, respectively) compared to ChatGPT (p < 0.05).
  • General dentists showed lower diagnostic accuracy (37.85%) compared to ChatGPT in most subgroups.

Abstract

Artificial intelligence (AI) is leading to a significant paradigm shift in medical imaging and diagnostic sciences. In particular, Chat Generative Pre-trained Transformer (ChatGPT) is finding increasing application in diagnostic processes due to its ability to generate clinical outcomes. This study aims to evaluate the diagnostic accuracy of ChatGPT for jawbone lesions and also to compare it with that of Oral and Maxillofacial Radiologists (OMFR), Oral and Maxillofacial Surgeons (OMFS), and general dentists. Thirty cases with jawbone lesions, for which clinical information, panoramic radiographs, and histopathological diagnoses were available, were selected. A questionnaire was prepared, including participants’ (OMFR, OMFS, and general dentists) demographic information, the cases’ clinical findings and panoramic radiographs, and distributed via electronic communication channels. The same cases were loaded into ChatGPT-4 and asked to generate a preliminary diagnosis. The data were statistically analyzed using the Wilcoxon Signed Rank, Mann–Whitney U, and Kruskal–Wallis tests at a significance level of p < 0.05. Overall, ChatGPT’s diagnostic accuracy was limited to 46.67%, while the OMFR (67.71%) and OMFS (58.96%) groups had statistically significantly higher success rates (p < 0.05) than ChatGPT. General dentists (37.85%) had lower or similar diagnostic accuracy compared to ChatGPT in most subgroups (gender, age, workplace, professional experience). ChatGPT demonstrated moderate diagnostic accuracy. While OMFR and OMFS participants had significantly higher accuracy rates than ChatGPT, ChatGPT generally outperformed general dentists. These results indicate that such AI systems cannot replace specialist clinicians but can provide valuable contributions as supportive tools that enhance diagnosis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Satan et al. (2026) studied this question.

synapsesocial.com/papers/6a0d5100f03e14405aa9d460https://doi.org/10.1186/s12903-026-08623-w
Ask AI
Helpful
Bookmark
Share
View Full Paper