PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026SHILAP Revista de lepidopterología3 citationsOpen Access

Evaluation of Large Language Models for Radiologists’ Support in Multidisciplinary Breast Cancer Teams: Comparative Study

View Full Paper
HJHong JiangCYChun YangWZWenbin Zhou

Key Points

  • This study evaluates the performance of large language models in assisting radiologists within breast cancer teams.
  • Developed 50 questions on radiological and breast cancer guidelines.
  • Posed questions to 9 popular large language models and clinical physicians.
  • Assessed performance based on accuracy, confidence, and consistency against standard guidelines.
  • Claude 3 Opus and ChatGPT-4 achieved high confidence scores (2.78 and 2.74).
  • ChatGPT-4o led in accuracy with a score of 2.92.
  • ChatGPT-4o mini excelled in clinical diagnostics with a score of 3.0, higher than all physician groups.
  • Significant score differences were noted only compared to fellow physicians.

Abstract

Background Artificial intelligence tools, particularly large language models (LLMs), have shown considerable potential across various domains. However, their performance in the diagnosis and treatment of breast cancer remains unknown. Objective This study aimed to evaluate the performance of LLMs in supporting radiologists within multidisciplinary breast cancer teams, with a focus on their roles in facilitating informed clinical decisions and enhancing patient care. Methods A set of 50 questions covering radiological and breast cancer guidelines was developed to assess breast cancer. These questions were posed to 9 popular LLMs and clinical physicians, with the expectation of receiving direct “Yes” or “No” answers along with supporting analysis. The performances of the 9 models, including ChatGPT-4.0, ChatGPT-4o, ChatGPT-4o mini, Claude 3 Opus, Claude 3.5 Sonnet, Gemini 1.5 Pro, Tongyi Qianwen 2.5, ChatGLM, and Ernie Bot 3.5, were evaluated against that of radiologists with varying experience levels (resident physicians, fellow physicians, and attending physicians). Responses were assessed for accuracy, confidence, and consistency based on alignment with the 2024 National Comprehensive Cancer Network Breast Cancer Guidelines and the 2013 American College of Radiology Breast Imaging-Reporting and Data System recommendations. Results Claude 3 Opus and ChatGPT-4 achieved the highest confidence scores of 2.78 and 2.74, respectively, while ChatGPT-4o led in accuracy with a score of 2.92. In terms of response consistency, Claude 3 Opus and Claude 3.5 Sonnet led the pack with scores of 3.0, closely followed by ChatGPT-4o, Gemini 1.5 Pro, and ChatGPT-4o mini, all recording impressive scores exceeding 2.9. ChatGPT-4o mini excelled in clinical diagnostics with a top score of 3.0 among all LLMs, and this score was also higher than all physician groups; however, no statistically significant differences were observed between it and any physician group (all P>.05). ChatGPT-4 also had a higher score than the physician groups but showed comparable statistical performance to them (P>.05). Across radiological diagnostics, clinical diagnosis, and overall performance, ChatGPT-4o mini and the Claude models achieved higher mean scores than all physician groups. However, these differences were statistically significant only when compared to fellow physicians (P.05). Among physician groups, attending physicians and resident physicians exhibited comparable high scores in radiological diagnostic performance, whereas fellow physicians scored somewhat lower, though the difference was not statistically significant (P>.05). Conclusions LLMs such as ChatGPT-4o and Claude 3 Opus showed potential in supporting multidisciplinary teams for breast cancer diagnostics and therapy. However, they cannot fully replicate the intricate decision-making processes honed through clinical experience, particularly in complex cases. This highlights the need for ongoing artificial intelligence refinement to ensure robust clinical applicability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jiang et al. (2026) studied this question.

synapsesocial.com/papers/698433c8f1d9ada3c1fb13bahttps://doi.org/10.2196/68182
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models2023 · 3,806 citations
  2. 2Artificial intelligence in cancer imaging: Clinical challenges and applications2019 · 1,874 citations
  3. 3Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries2024 · 25,401 citations
  4. 4Evaluation of large language models in breast cancer clinical scenarios: a comparative analysis based on ChatGPT-3.5, ChatGPT-4.0, and Claude22024 · 90 citations
  5. 5The Impact of Reasoning Step Length on Large Language Models2024 · 3 citations