PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 5, 2025BMC Emergency Medicine0 citationsOpen Access

Performance of ChatGPT, Gemini and DeepSeek for non-critical triage support using real-world conversations in emergency department

View Full Paper
SLSukyo LeeSJSumin JungJPJong Park

Key Points

  • Gemini 2.5 flash achieved the highest triage accuracy at 73.8%, demonstrating strong potential for AI in emergency care.
  • A total of 1,057 triage conversations were analyzed, revealing significant variations in model performance across different LLMs.
  • Using both zero-shot and few-shot prompting improved outcomes, highlighting the flexibility and adaptability of LLMs in clinical situations.
  • The findings support the integration of LLMs for non-critical triage, benefiting patient care in diverse clinical environments.

Abstract

Abstract Background Timely and accurate triage is crucial for the emergency department (ED) care. Recently, there has been growing interest in applying large language models (LLMs) to support triage decision-making. However, most existing studies have evaluated these models using simulated scenarios rather than real-world clinical cases. Therefore, we evaluated the performance of multiple commercial LLMs for non-critical triage support in ED using real-world clinical conversations. Methods We retrospectively analyzed real-world triage conversations prospectively collected from three tertiary hospitals in South Korea. Multiple commercial LLMs—including OpenAI GPT-4o, GPT-4.1, O3, Google Gemini 2.0 flash, Gemini 2.5 flash, Gemini 2.5 pro, DeepSeek V3, and DeepSeek R1—were evaluated for the accuracy in triaging patient urgency based solely on unsummarized dialogue. The Korean Triage and Acuity Scale (KTAS) assigned by triage nurses was used as the gold standard for evaluating the LLM classifications. Model performance was assessed under both a zero-shot prompting condition and a few-shot prompting condition that included representative examples. Results A total of 1,057 triage cases were included in the analysis. Among the models, Gemini 2.5 flash achieved the highest accuracy (73.8%), specificity (88.9%), and PPV (94.0%). Gemini 2.5 pro demonstrated the highest sensitivity (90.9%) and F1-score (82.4%), though with lower specificity (23.3%). GPT-4.1 also showed balanced high accuracy (70.6%) and sensitivity (81.3%) with practical response times (1.79s). Performance varied widely between models and even between different versions from the same vendor. With few-shot prompting, most models showed further improvements in accuracy and F1-score. Conclusions LLMs can accurately triage ED patient urgency using real-world clinical conversations. Several models demonstrated both high sensitivity and acceptable response times, supporting the feasibility of LLM in non-critical triage support tools in diverse clinical environments. These findings apply to non-critical patients (KTAS 3–5), and further research should address integration with objective clinical data and real-time workflow.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lee et al. (2025) studied this question.

synapsesocial.com/papers/68bb4d196d6d5674bcd00b92https://doi.org/10.1186/s12873-025-01337-2
Ask AI
Helpful
Bookmark
Share
View Full Paper