PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026iScience0 citationsOpen Access

Comparison of the performance of large language models in answering patient questions related to cataract

View Full Paper
QHQing HeJSJiayi ShiXLXinyi Liu

Key Points

Key points are not available for this paper at this time.

Abstract

This study evaluated the performance of four popular large-scale language models (ChatGPT o3-mini, Gemini 2.0 pro experimental, Deep Seek Thinking R1, and Kimi Thinking K1.5) in addressing frequently asked patient questions about cataracts and cataract surgery in Chinese. DeepSeek Thinking R1 performed comparably to Gemini 2.0 pro experimental in accuracy, while outperforming both ChatGPT o3-mini and Kimi Thinking K1.5. In terms of completeness and consistency, DeepSeek Thinking R1 showed superior performance over the other three LLMs. Regarding legibility and safety, DeepSeek Thinking R1, Gemini 2.0 pro experimental, and ChatGPT o3-mini exhibited comparable results, all performing better than Kimi Thinking K1.5. Deep Seek Thinking R1 demonstrated the strongest overall performance among the four LLMs in this comparative evaluation. The modern LLMs are promising tools for public education in ophthalmology while human oversight is still required.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

He et al. (2026) studied this question.

synapsesocial.com/papers/6a0889d0df3db87398109e9chttps://doi.org/10.1016/j.isci.2026.115002
Ask AI
Helpful
Bookmark
Share
View Full Paper