PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 2026Journal of the American Medical Informatics Association0 citations

Development of BERT-based large language models for emergency department triage using real-world conversations

View Full Paper
SLSukyo LeeSJSumin JungJPJong-Hak Park

Key Points

  • The research aimed to develop BERT-based language models trained on real-world conversations to enhance triage accuracy in emergency departments.
  • Developed two BERT-based models using a dataset of anonymized triage-level conversations.
  • One model tokenized entire conversations, the other used a hierarchical structure with sentence-level tokenization.
  • Evaluated model performance through accuracy, precision, recall, and F1-score.
  • Compared performance against ChatGPT and ClinicalBERT, assessing explainability using SHAP.
  • The hierarchical model achieved an accuracy of 75.94%, significantly surpassing ChatGPT's 56.68% and ClinicalBERT's 69.42%.
  • Best model recorded a recall of 0.9610 for urgent cases, outperforming ChatGPT's 0.5352.
  • SHAP analysis confirmed the model focused on clinically relevant cues corresponding to KTAS criteria.

Abstract

Abstract Objectives Accurate triage in emergency departments (ED) is critical for appropriate resource allocation. While artificial intelligence (AI) has been explored for triage, prior models relied on summarized clinical scenarios. We aimed to develop and evaluate large language models (LLMs) trained on real-world clinical conversations to classify patient urgency. Materials and Methods We used a nationally curated dataset of anonymized triage-level conversations from 3 tertiary Korean hospitals. Two BERT-based models were developed to classify urgency per the Korean Triage and Acuity Scale (KTAS) into urgent (KTAS 3) or non-urgent (KTAS 4-5). One model tokenized the entire conversation, while the other applied a hierarchical structure with sentence-level tokenization and speaker-role embeddings. Performance metrics included accuracy, precision, recall, and F1-score. We compared our models against ChatGPT GPT-4o and ClinicalBERT, and assessed explainability using SHapley Additive exPlanations (SHAP). Results A total of 5244 clinical conversations, 1057 triage-level dialogues were used, with 950 for training and 107 for testing. Our model with hierarchical structure achieved accuracies of 75.94%, significantly outperforming ChatGPT (56.68%) or fine-tuned ClinicalBERT (69.42%). For urgent cases, the best model achieved a recall of 0.9610, outperforming ChatGPT (0.5352). SHapley Additive exPlanations analysis confirmed that our model focused on clinically relevant cues aligned with KTAS criteria. Conclusion BERT-based LLMs trained on real-world ED conversations significantly outperform general-purpose models like ChatGPT in triage accuracy. This approach demonstrates the potential for enhancing clinical decision support with interpretable and efficient AI.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lee et al. (2026) studied this question.

synapsesocial.com/papers/698828770fc35cd7a8847f78https://doi.org/10.1093/jamia/ocag007
Ask AI
Helpful
Bookmark
Share
View Full Paper