PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026International Journal of Surgery1 citationsOpen Access

A comparative evaluation of large language models for simplifying prostate cancer pathology reports: ChatGPT and Gemini

View Full Paper
HZHaoyang ZengYYYangguang YuanXWXiang Wu

Key Points

  • This study aims to evaluate and compare the effectiveness of various AI models in simplifying prostate cancer pathology reports.
  • Evaluated three versions of ChatGPT and Gemini on pathology reports from 228 prostate cancer patients.
  • Data was split into internal and external cohorts, using specific prompts for text generation.
  • Outputs were assessed based on human scoring, readability scores, and BERT-based semantic similarity scores.
  • GPT-4o (Few-Shot) received the highest scores for accuracy and comprehensiveness from pathologists.
  • Gemini showed the best understandability ratings across all groups.
  • Mean Reading Grade Level scores varied, but GPT-4o Few-Shot performed best overall.

Abstract

Objectives: To evaluate the application value of three ChatGPT versions and Gemini in pathology report simplification tasks for prostate cancer. Methods: This retrospective study assessed GPT-3.5, GPT-4.0, GPT-4o, and Gemini on pathology reports from 228 prostate cancer patients across two institutions. Data were split into internal (center 1, n = 171) and external (center 2, n = 57) cohorts. Using specific prompts, models generated simplified texts. The evaluation of outputs included three main dimensions: (1) human scoring by patients, clinicians, and pathologists; (2) readability scores; and (3) BERT-based semantic similarity scores. Statistical comparisons employed paired t -tests or Wilcoxon signed-rank tests. Statistical consistency between raters was assessed using squared weighted kappa, intraclass correlation coefficient(3,1), and percent agreement, with 95% confidence intervals calculated for all metrics. Results: GPT-4o (Few-Shot) achieved the highest accuracy and comprehensiveness scores from pathologists, while Gemini demonstrated the best understandability. Patient and clinician understandability ratings were consistently high across models. Mean Reading Grade Level scores varied between internal and external datasets, with GPT-4o Few-Shot performing best overall. BERT-based semantic similarity scores demonstrated distinct trends across models, reflecting differences in text simplification strategies. Conclusion: LLMs adopt distinct trade-off strategies between simplifying pathology reports and preserving their structure and logic, influenced by prompt design and textual style. Their application shows potential to enhance patient comprehension and clinical communication. Future work should focus on domain-specific fine-tuning to ensure safe and reliable clinical integration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zeng et al. (2026) studied this question.

synapsesocial.com/papers/6984359ef1d9ada3c1fb49c3https://doi.org/10.1097/js9.0000000000004454
Ask AI
Helpful
Bookmark
Share
View Full Paper