Background/Objectives: To evaluate the numerical agreement and variability between large language model (LLM)-generated surgical dosage recommendations and expert surgical planning in strabismus surgery under controlled conditions. Methods: This retrospective, single-center study included patients who underwent horizontal strabismus surgery between January 2023 and May 2025 and achieved successful postoperative alignment (≤±10 prism diopters at 3 months). Standardized preoperative clinical data were provided to three LLMs (ChatGPT-5.2, Gemini 3.0 Pro, and DeepSeek v3.2), each prompted to generate surgical dosage recommendations in millimeters. Model outputs were compared with the surgeon’s preoperative plans. Agreement and variability were assessed using intraclass correlation coefficients (ICC) and Bland–Altman analyses. Results: A total of 68 patients were included (44.1% female; mean age 11.7 ± 10.5 years). Agreement between LLM-generated and surgeon-determined surgical dosages was heterogeneous across models and surgical subtypes. Moderate ICC values were observed in selected subgroups; however, confidence intervals were frequently wide and, in several instances, crossed zero, indicating limited statistical reliability. Bland–Altman analyses demonstrated wide limits of agreement, often approaching or exceeding ±1 mm. Agreement tended to be higher in smaller deviation angles (≤35 prism diopters) and lower in larger deviations and more complex surgical scenarios. Conclusions: LLM-generated surgical dosage recommendations demonstrated variable and context-dependent agreement with expert planning, with substantial variability and limited reliability across subgroups. These findings indicate that current LLM-based systems do not achieve the level of precision required for quantitative surgical decision-making and should be considered investigational and not suitable for independent clinical use. This study represents a variability and calibration analysis rather than an evaluation of clinical accuracy or decision-making performance.
Koçkar et al. (Tue,) studied this question.