PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 4, 2026JCO Clinical Cancer Informatics1 citations

Generative Artificial Intelligence for Medical Summarization in Prostate Cancer: Comparative Evaluation by Physicians and Patient Advocates—A Pilot Study

View Full Paper
CRC. RaynaudLDLoris DematiniJBJean‐Emmanuel Bibault

Key Points

  • This research aims to evaluate the effectiveness of different generative AI models in summarizing medical information on prostate cancer from both physician and patient perspectives.
  • Conducted a survey-based pilot study involving four generative AI models: Llama 3, Mistral Large 2, Gemma 2B, and Consensus.
  • Selected seven recent prostate cancer abstracts for summarization and translation tasks.
  • Physicians and patients evaluated the summaries using structured Likert-scale questionnaires for various quality criteria.
  • Performed descriptive statistical analyses on the evaluation results.
  • Consensus model received the highest ratings from physicians across all criteria including completeness and accuracy.
  • Gemma 2B was favored by patients for conciseness and clarity, while Consensus excelled in organization.
  • Overall positive evaluations were noted for Llama 3 and Mistral Large 2, although they received fewer strongly agree ratings.
  • Marked differences in perceived accuracy and clarity were observed across the different AI models.

Abstract

PURPOSE The exponential growth of scientific publications presents increasing challenges for clinicians and patients seeking to access up-to-date medical information. Language models (LMs) have emerged as powerful tools for generating and summarizing scientific content, but their performance in oncology remains insufficiently characterized from both professional and patient perspectives. MATERIALS AND METHODS We conducted a prospective, survey-based pilot study evaluating four LMs: Llama 3, Mistral Large 2, Gemma 2B, and Consensus, applied to the summarization and French translation of seven recent prostate cancer (PCa) abstracts. Each model received a standardized prompt to generate the summary of each abstract. Physicians (medical and radiation oncologists, urologists) and patients treated for PCa independently assessed the outputs using structured Likert-scale questionnaires covering qualitative criteria such as accuracy, usefulness, organization, and comprehensibility. Descriptive statistical analyses were then performed to characterize the distribution of responses across evaluation items. RESULTS A total of 40 respondents (14 physicians, 26 patients) provided 280 individual evaluations. Across physicians, consensus received the highest proportion of strongly agree ratings for all criteria, including completeness, accuracy, currency, organization, and usefulness. Only two physicians evaluated this model. Among patients, Gemma 2B achieved the highest strongly agree ratings for conciseness and comprehensibility, whereas Consensus obtained the highest score for organization; both models were similarly rated for usefulness. When considering overall positive evaluations (agree or strongly agree), Llama 3 and Mistral Large 2 performed well across groups but generated fewer strongly agree responses. Descriptive analyses demonstrated clear differences in perceived accuracy, completeness, and clarity across model architectures. CONCLUSION Perceived quality varied across models and user groups. Consensus was preferred by physicians, whereas patients more often favored Gemma. These differences underscore the importance of selecting models aligned with specific clinical communication tasks when deploying generative artificial intelligence in oncology.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Raynaud et al. (2026) studied this question.

synapsesocial.com/papers/69d0af52659487ece0fa5321https://doi.org/10.1200/cci-25-00316
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Toward Automating the Summarization of Cancer Pathology Reports Using Large Language Models to Improve Clinical Usability2026 · 1 citations
  2. 2Patient-Centered Summarization Framework for AI Clinical Summarization: Mixed Methods Study2026
  3. 3Benchmarking next-generation large language models (LLMS): Evaluation of patient-oriented prostate cancer guidance.2026
  4. 4Patient friendly summaries of oncology consultations generated by large language models - A pilot study of patient and provider satisfaction2025
  5. 5Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI (Preprint)2025