GPT-4o and Gemini 2.0 Flash classified ASA physical status from synthetic scenarios with almost perfect agreement against expert consensus (κw=0.936 and 0.933, respectively; p=1.0 for accuracy difference).
Do large language models (GPT-4o and Gemini 2.0 Flash) accurately classify ASA physical status from synthetic preoperative scenarios compared to expert consensus?
GPT-4o and Gemini 2.0 Flash can classify ASA physical status from synthetic preoperative scenarios with almost perfect agreement against expert consensus, highlighting their potential to standardize preoperative risk stratification.
Absolute Event Rate: 0.936% vs 0.933%
p-value: p=1.0
ABSTRACTBackground: The ASA Physical Status (PS) classification is the most widely usedpreoperative risk stratification tool in anesthesiology, yet its application is characterized bysubstantial inter-rater variability. Large language models (LLMs) may offer a means tostandardize this inherently subjective assessment. This study evaluates GPT-4o andGemini 2.0 Flash on expert-validated synthetic preoperative scenarios, a design thatenables rigorous LLM benchmarking without the ethical and regulatory constraints ofpatient data use.Methods: We developed 487 synthetic preoperative scenarios spanning the fullASA I–IV spectrum across multiple surgical specialties. A panel of three board-certifiedanesthesiologists independently assigned ASA class and reached consensus gold-standardlabels through a structured adjudication protocol. GPT-4o (gpt-4o-2024-11-20) andGemini 2.0 Flash (gemini-2.0-flash-001) were queried three times per scenario usingstandardized zero-shot Turkish-language prompts (temperature=0). The primary outcomewas quadratic weighted kappa (κw).Results: The 487 scenarios comprised ASA I (26.5%, n=129), II (23.6%, n=115),III (41.9%, n=204), and IV (8.0%, n=39). Expert panel inter-rater reliability was κw=0.920(95%CI 0.896–0.945). GPT-4o achieved κw=0.936 (95%CI 0.914–0.958) and Gemini 2.0κw=0.933 (95%CI 0.911–0.955), both exceeding the 'almost perfect' threshold (κw>0.80).Overall accuracy was 91.8% (GPT-4o) and 91.6% (Gemini); models did not differsignificantly (McNemar p=1.0). For ASA≥III high-risk identification, GPT-4o achievedsensitivity 95.9% and specificity 96.7%; Gemini achieved 96.7% and 98.0%, respectively.Conclusions: GPT-4o and Gemini 2.0 Flash classify ASA physical status fromsynthetic preoperative scenarios with almost perfect agreement against expert consensus,matching the expert panel's own inter-rater reliability. Prospective clinical validation is thelogical next step.
Karaoğlu et al. (Sat,) conducted a other in ASA Physical Status classification (n=487). GPT-4o vs. Gemini 2.0 Flash was evaluated on quadratic weighted kappa (κw) (95% CI 0.914-0.958, p=1.0). GPT-4o and Gemini 2.0 Flash classified ASA physical status from synthetic scenarios with almost perfect agreement against expert consensus (κw=0.936 and 0.933, respectively; p=1.0 for accuracy difference).