PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 28, 2026BMC Medical Informatics and Decision Making0 citationsOpen Access

Application of large language models in clinical decision-making for dry eye disease

HCHaoqiang CuiYYYuyang YangTLTaichen Lai

Key Points

  • Evaluate the performance of large language models in clinical decision-making for dry eye disease compared to junior ophthalmologists.
  • Analyzed 100 standardized cases of dry eye disease over a two-year period.
  • Four large language models were tested for diagnostic and treatment decision performance.
  • Evaluated using Global Quality Score and clinical safety ratings by senior ophthalmologists.
  • Large language models achieved treatment necessity accuracy between 96-99%, comparable to junior ophthalmologists’ 97%.
  • DeepSeek-V3 demonstrated the highest classification accuracy of 92% with strong expert agreement (kappa = 0.80).
  • LLMs had a significantly shorter response time (16-36 s) compared to junior ophthalmologists (315 ± 57 s).

Abstract

To systematically evaluate the diagnostic, classification, and treatment decision-making performance of large language models (LLMs) in dry eye disease (DED) and compare their performance with that of junior ophthalmologists to assess their feasibility as clinical support tools. One hundred standardized DED cases from Fuzhou University Affiliated Provincial Hospital (Aug 2023–Aug 2025) were analyzed. Four LLMs (ChatGPT-4o, DeepSeek-V3, Gemini-2.5-Pro, and ERNIE Bot-4.5-turbo) underwent an initial 20-item test with a 95% accuracy threshold. Only models meeting this criterion advanced to case-based evaluation. Diagnostic and therapeutic outputs were rated by senior ophthalmologists using the Global Quality Score (GQS) and a clinical safety score. Meanwhile, the response times of both the LLMs and the junior ophthalmologists were recorded to evaluate their respective efficiency. ChatGPT-4o, DeepSeek-V3, and Gemini-2.5-Pro met the screening threshold. In determining treatment necessity, the accuracies of the LLMs (96–99%) were comparable to those of junior ophthalmologists (97%). For DED classification, DeepSeek-V3 achieved the highest accuracy (92%) and agreement with experts (kappa = 0.80), outperforming ChatGPT-4o and junior ophthalmologists (71%, kappa ≈ 0.53; p < 0.01). Gemini-2.5-Pro showed strong performance (accuracy 89%), the highest GQS (4.87 ± 0.37), and the best safety rating, with 89% of its outputs rated as “good”. The decision efficiency of the LLMs was significantly higher (p ≤ 0.001), with response times (16–36 s) much shorter than those of junior ophthalmologists (315 ± 57 s). Gemini-2.5-Pro and DeepSeek-V3 demonstrated high accuracy, safety, and efficiency in case-based DED management, showing strong potential as auxiliary tools to enhance clinical decision-making and support less experienced clinicians. Not applicable. This study is a retrospective observational study and does not involve any intervention requiring trial registration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cui et al. (2026) studied this question.

synapsesocial.com/papers/6a17dcbb3fad632b0f9d9751https://doi.org/10.1186/s12911-026-03586-y
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Comparative Evaluation of Deep-Reasoning Large Language Models for Ophthalmic Emergencies2026 · 1 citations
  2. 2Evaluating the competence of large language models in ophthalmology clinical practice: a multi-scenario quantitative study2025 · 1 citations
  3. 3Application of Large Language Models in Complex Clinical Cases: Cross-Sectional Evaluation Study2025 · 5 citations
  4. 4Evaluation of the Performance of 3 Large Language Models in Clinical Decision Support: A Comparative Study Based on Actual Cases (Preprint)2024
  5. 5Large Language Models Use in Dry Eye Disease: Perplexity AI versus ChatGPT42025