PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 29, 2026Medical Science Monitor0 citationsOpen Access

Preliminary Evaluation of Large Language Models in Kennedy Classification and Removable Partial Denture Planning: An In Silico Study

EEEbru SÜMER EKİNYDYalçın DeğerBYBerivan Dündar Yılmaz

Key Points

  • This study aims to evaluate the accuracy of large language models in Kennedy classification and removable partial denture planning.
  • Twenty-six partially edentulous cases were evaluated according to the FDI tooth numbering system.
  • Questions were submitted to four large language models without training: ChatGPT-5.2, Claude Sonnet 4.6, Gemini 3 Flash, and Perplexity.
  • Responses were assessed by two independent prosthodontists using predefined criteria for accuracy.
  • Significant differences were found among models in Kennedy classification accuracy (P<0.001, moderate effect size, Kendall's W=0.402).
  • Gemini 3 Flash scored highest in Kennedy classification, outpacing ChatGPT-5.2, Perplexity, and Claude Sonnet 4.6.
  • In removable partial denture planning, Gemini 3 Flash again achieved the highest performance (P<0.001, Kendall's W=0.424).

Abstract

Background:This study evaluated text-based large language models for Kennedy classification and removable partial denture planning in partially edentulous cases, where comparative evidence on accuracy remains limited. Material/Methods:Twenty-six partially edentulous cases, defined according to the Fdration Dentaire Internationale tooth numbering system and classified as Kennedy Classes I to IV, were included.Case scenarios were constructed by an experienced prosthodontist.Questions were independently submitted to 4 large language models without prior training or prompting: ChatGPT-5.2,Claude Sonnet 4.6, Gemini 3 Flash, and Perplexity.Responses were evaluated by 2 independent prosthodontists (different from the case developer) -blinded to the identity of each model -using predefined criteria for Kennedy classification accuracy and prosthetic planning consistency.Statistical analyses included intergroup comparisons and effect size estimation. Results:Statistically significant differences were identified among the models in both tasks (P<0.001).For Kennedy classification, the effect size was moderate (Kendall's W=0.402).Gemini 3 Flash achieved the highest mean score, followed by ChatGPT-5.2,Perplexity, and Claude Sonnet 4.6.Similarly, significant differences were observed in removable partial denture planning performance (Kendall's W=0.424); Gemini 3 Flash scored significantly higher than the other models (P0.001). Conclusions:Large language models showed variable performance under zero-shot conditions.Although Gemini 3 Flash achieved higher scores, the moderate effect sizes warrant cautious interpretation.These models may serve as adjunctive decision-support tools in prosthodontics but cannot replace clinical judgment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

EKİN et al. (2026) studied this question.

synapsesocial.com/papers/69f19f74edf4b468248063d4https://doi.org/10.12659/msm.953353
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Effectiveness of Various General large language models in Clinical Consensus and Case Analysis in Dental Implantology: A Comparative Study2024 · 1 citations
  2. 2A comparative evaluation of large language models in diagnosis and treatment planning in restorative dentistry2026
  3. 3Accuracy and Completeness of Contemporary Large Language Models in Prosthodontics: An Expert-Based Comparative Study2026
  4. 4Clinical Relevance of Large Language Models in Endodontics: Diagnostic Appropriateness Based on 50 Simulated Case Scenarios2025
  5. 5Clinical reliability and consistency of large language models in prosthodontics: an expert evaluation of ChatGPT and Gemini2026