PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
Synapse
⌘+K
Synapse
May 10, 2026Scientific ReportsOpen Access

Benchmarking large language models on persian surgical subspecialty board examinations: a comparative study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash

View Full Paper
Ask AI
Bookmark
Share

Authors

SSShahab SheikhalishahiFRFarzad RafieiSHSeyed Masoud Hosseini

Discussion

Loading...

Member takes

Overview

Comparative study evaluates accuracy of language models on surgical questions, indicating areas for improvement.

Key Points

  • This research assesses how well different large language models perform on Persian surgical board examination questions.
  • Evaluated three models: ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash on 532 questions.
  • Questions spanned five surgical domains: Pediatric, Cardiovascular, Vascular and Endovascular, Thoracic, and Plastic & Reconstructive Surgery.
  • Measured accuracy, agreement with official keys, and the impact of question length on performance.
  • Gemini 2.5 Flash achieved 73.9% accuracy and ChatGPT-5 73.3%, both higher than ChatGPT-4o at 68.2%.
  • Substantial agreement with official keys for Gemini 2.5 Flash (κ = 0.651) and ChatGPT-5 (κ = 0.642).
  • All models performed poorly on surgical technique questions compared to clinical scenarios, with question length negatively impacting ChatGPT-4o's performance.

Cite This Study

Sheikhalishahi et al. (2026) studied this question.

synapsesocial.com/papers/6a0021fec8f74e3340f9cf7chttps://doi.org/10.1038/s41598-026-51934-9
View Full Paper
Ask AI
Bookmark
Share