PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 4, 2026Journal of Clinical Oncology0 citations

Using large language models for grading CTCAE toxicity after radiation therapy for prostate cancer.

View Full Paper
RWRenthony WilsonFMFederico MastroleoMOMariana Borras Osorio

Key Points

  • The study aims to assess how well large language models can extract and grade CTCAE toxicities from clinical notes of prostate cancer patients.
  • Analyzed clinical notes and patient-reported outcomes from 55 prostate cancer patients undergoing proton therapy.
  • Developed a prompt to extract and grade toxicity using large language models (OpenAI GPT-4o and Gemini 2.0 Flash).
  • Conducted two phases of evaluation to refine performance metrics against a ground truth of CTCAE graded toxicities.
  • Created a conservative stepwise ensemble model to combine strengths of both LLMs.
  • Achieved binary accuracy of 0.954 with a sensitivity of 0.961 and specificity of 0.953 using the ensemble model.
  • The F1 score was 0.836, indicating a good balance between precision and recall.
  • Gemini 2.0 Flash and OpenAI GPT-4o demonstrated strong metrics in both phases of evaluation.

Abstract

357 Background: Prostate cancer (PCa) is the most incident cancer in adult males in the United States, and the second highest in the world. CTCAE is the standardized method to grade the severity of treatment-related adverse events (AEs), but are tedious to collect and subject to inter-observer variability. This study aimed to evaluate the performance of LLMs in automatically extracting and grading CTCAE toxicities from clinical notes and patient-reported outcomes of PCa patients from a clinical trial. Methods: A total of 55 patients from NCT02874014 trial, undergoing proton therapy (PT) for localized PCa, were included. The ground truth of CTCAE graded toxicities was obtained from the trial’s results. Clinical notes within 3 days, and the most recent patient questionnaires within 90 days of toxicity assessment were retrieved from the electronic health record. A comprehensive prompt was developed to extract and CTCAE grade toxicities from the retrieved records. LLMs used for this analysis were OpenAI GPT-4o and Gemini 2.0 Flash. Phase 1 evaluated both LLMs' performance against the ground truth. To ensure the metrics reflected the models’ performance based only on the information available in the text provided, cases where both LLMs contradicted the ground truth were manually reviewed. Phase 2 recalculated LLMs' performance metrics using this adapted ground truth. A conservative stepwise ensemble model was developed to leverage the complementary strengths of both LLMs. Results: In the conservative stepwise ensemble model, Gemini 2.0 Flash served as the initial screening platform to identify AEs in step 1. All cases with identified toxicity events were then reviewed and confirmed using OpenAI GPT-4o in step 2. This effectively reduced false positive rates while preserving high sensitivity. The ensemble model achieved strong performance metrics: binary accuracy (0.954), sensitivity (0.961), specificity (0.953), F1 score (0.836), and grade accuracy (0.932). Conclusions: This study demonstrates the feasibility of the use of LLMs for extraction and CTCAE toxicity grading in PCa patients treated with RT. The ensemble model achieved strong performance metrics. Further validation is needed for other cancer diagnosis and treatment modalities. Phase Model Binary Accuracy Precision Sensitivity Specificity F1 Score Grade Accuracy 1 Gemini 0.946 0.672 1.000 0.939 0.804 0.919 1 OpenAI 0.940 0.785 0.637 0.978 0.703 0.895 2 Gemini 0.956 0.735 1.000 0.950 0.848 0.927 2 OpenAI 0.954 0.917 0.679 0.991 0.780 0.903

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wilson et al. (2026) studied this question.

synapsesocial.com/papers/69a7ccf7d48f933b5eed8dc8https://doi.org/10.1200/jco.2026.44.7_suppl.357
Ask AI
Helpful
Bookmark
Share
View Full Paper