Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
April 17, 2026JA Clinical ReportsOpen Access

Evaluation of the reliability of large language models for ASA-PS classification in cardiovascular surgery: a pilot study

View Full Paper
Ask AI
Bookmark
Share

Authors

KIKeisuke IwabuTJTakashi JuriSTShogo Tsujikawa

Discussion

Loading...

Member takes

Overview

Pilot study evaluates LLM reliability for ASA-PS classification in cardiovascular surgery, indicating potential as training adjuncts.

Key Points

  • This study aims to assess the reliability of large language models for ASA-PS classification in cardiovascular surgery.
  • Rated 32 anonymized cases by two residents and two board-certified cardiovascular anesthesiologists.
  • Evaluated four large language model modes including ChatGPT and Gemini.
  • Utilized zero-shot evaluation for model assessments.
  • Calculated overall agreement using intraclass correlation coefficients.
  • Moderate overall agreement among evaluators (ICC 0.49–0.52).
  • Good agreement between LLMs and specialists (ICC 0.61–0.65).
  • Exact-match rates were 42.2% for residents and 59.4–75.0% for LLMs.
  • Classifications outside expert ranges were rare (0–3.1%).

Cite This Study

Iwabu et al. (2026) studied this question.

synapsesocial.com/papers/69e1cf985cdc762e9d85886chttps://doi.org/10.1186/s40981-026-00858-4
View Full Paper
Ask AI
Bookmark
Share