PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026Veterinary Record0 citationsOpen Access

Assessing the performance of large language models when used to determine ASA status of cats and dogs and generate anaesthetic protocols

View Full Paper
SOSıtkıcan OkurYAYasemin AkçoraDODamla T. Okur

Key Points

  • This study aims to evaluate how well large language models can determine ASA status and generate anesthetic protocols in veterinary medicine.
  • Retrospective analysis of 225 feline and canine cases with ASA classifications from a veterinary hospital.
  • Three large language models (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Pro) assessed ASA classifications and generated protocols.
  • Statistical analyses included Friedman and Bonferroni-adjusted Wilcoxon tests and inter-panelist reliability measures.
  • ChatGPT-5 achieved the highest ASA classification accuracy at 53.3%, compared to ChatGPT-4o (46.7%) and Gemini 2.5 Pro (30.7%).
  • Performance was strongest for ASA classifications 3‒5, with significant misclassification observed in ASA 1 cases.
  • ChatGPT-5 generated the most clinically adequate anesthetic protocols, surpassing the other models.

Abstract

BACKGROUND: Large language models (LLMs) are emerging as decision-support tools in human medicine; however, their evaluation in veterinary anaesthesiology remains limited. METHODS: We retrospectively analysed 225 anonymised feline and canine cases (American Society of Anaesthesiologists ASA classifications 1‒5) from Atatürk University Veterinary Hospital. ChatGPT-4o, ChatGPT-5 and Gemini 2.5 Pro independently assigned ASA classifications and generated anaesthetic protocols using standardised prompts. Protocol adequacy was evaluated for all cases, regardless of ASA classification agreement, by two experienced veterinary anaesthesiologists using a four-point scale. Statistical analyses included Friedman and Bonferroni-adjusted Wilcoxon tests, effect sizes and inter-panelist reliability (assessed by quadratic-weighted Cohen's kappa and intraclass correlation coefficient). RESULTS: ChatGPT-5 achieved the highest ASA classification accuracy (53.3%), followed by ChatGPT-4o (46.7%) and Gemini 2.5 Pro (30.7%). The performance was strongest for ASA 3‒5, whereas ASA 1 cases were frequently misclassified, mainly due to ASA overestimation. ChatGPT-5 generated the most clinically sufficient anaesthetic protocols, outperforming the other models. LIMITATIONS: The retrospective, single-centre design and inclusion of only feline and canine cases may limit generalisability. CONCLUSIONS: LLMs can generate clinically relevant ASA classifications and anaesthetic protocols in veterinary anaesthesiology, although performance varies across models. However, expert oversight remains essential.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Okur et al. (2026) studied this question.

synapsesocial.com/papers/6a0566fba550a87e60a1efaehttps://doi.org/10.1002/vetr.70741
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of large language models on veterinary undergraduate multiple-choice examinations: a comparative evaluation2025
  2. 2Assessing the Role of Large Language Models in Veterinary Dentistry Client Communication2026
  3. 3Performance of Large Language Models in ASA Physical Status Classification Using Expert-Validated Synthetic Preoperative Scenarios: A Comparative Study of GPT-4o and Gemini 2.02026
  4. 4Performance of large language models versus clinicians and novices in veterinary theriogenology decision support2026 · 1 citations
  5. 5Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures2026