Benchmarking large language models on persian surgical subspecialty board examinations: a comparative study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash
Comparative study evaluates accuracy of language models on surgical questions, indicating areas for improvement.
Key Points
This research assesses how well different large language models perform on Persian surgical board examination questions.
Evaluated three models: ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash on 532 questions.
Questions spanned five surgical domains: Pediatric, Cardiovascular, Vascular and Endovascular, Thoracic, and Plastic & Reconstructive Surgery.
Measured accuracy, agreement with official keys, and the impact of question length on performance.
Gemini 2.5 Flash achieved 73.9% accuracy and ChatGPT-5 73.3%, both higher than ChatGPT-4o at 68.2%.
Substantial agreement with official keys for Gemini 2.5 Flash (κ = 0.651) and ChatGPT-5 (κ = 0.642).
All models performed poorly on surgical technique questions compared to clinical scenarios, with question length negatively impacting ChatGPT-4o's performance.
Cite This Study
Sheikhalishahi et al. (2026) studied this question.