PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 30, 2026iNew Medicine1 citationsOpen Access

Medical Reasoning With Large Language Models: A Systematic Review and Evaluation

View Full Paper
XRXiaohan RenCFChenxiao FanWMWenyin Ma

Key Points

Key points are not available for this paper at this time.

Abstract

ABSTRACT Large language models (LLMs) have achieved strong performance on medical exam–style tasks, motivating growing interest in their deployment in real‐world clinical settings. However, clinical decision‐making is inherently safety‐critical, context‐dependent, and conducted under evolving evidence. In such situations, reliable LLM performance depends not on factual recall alone but on robust medical reasoning. In this work, we present a comprehensive review of medical reasoning with LLMs. Grounded in cognitive theories of clinical reasoning, we conceptualize medical reasoning as an iterative process of abduction, deduction, and induction, and we organize existing methods into seven major technical routes spanning training‐based and training‐free approaches. We further conduct a unified cross‐benchmark evaluation of representative medical reasoning models under a consistent experimental setting, enabling a more systematic and comparable assessment of the empirical impact of existing methods. To better assess clinically grounded, decision‐oriented reasoning, we introduce MR‐Bench, a benchmark derived from real‐world hospital data. Evaluations on MR‐Bench expose a pronounced gap between exam‐level performance and accuracy on authentic clinical decision tasks. Overall, this survey provides a unified view of existing medical reasoning methods, benchmarks, and evaluation practices and highlights key gaps between current model performance and the requirements of real‐world clinical reasoning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ren et al. (2026) studied this question.

synapsesocial.com/papers/6a1c1be3c97d63156a5f537dhttps://doi.org/10.1002/inm3.70056
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Humanity's Last Exam2025 · 18 citations
  2. 2Data-Driven Prediction of Drug Effects and Interactions2012 · 963 citations
  3. 3Dual Process Theory for Large Language Models: An overview of using Psychology to address hallucination and reliability issues2023 · 19 citations
  4. 4Large language models in real-world clinical workflows: a systematic review of applications and implementation2025 · 68 citations
  5. 5Abductive reasoning and clinical assessment1997 · 26 citations