PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 8, 2026Frontiers in Oncology0 citationsOpen Access

Accuracy of large language models in head and neck cancers: a comparative analysis of ChatGPT and Gemini in TNM staging and clinical decision support

View Full Paper
DBDeniz BaklaciHIHuseyin IsikDEDuygu Erdem

Key Points

Key points are not available for this paper at this time.

Abstract

Objective Large Language Models (LLMs) are increasingly integrated into oncological workflows. This study evaluated and compared the performance of ChatGPT-4o and Gemini 1.5 Pro in TNM staging and treatment planning for head and neck malignancies across diverse anatomical sites. Materials and Methods A retrospective analysis was performed on 180 patients with head and neck cancer. Clinical data were processed through structured prompts to simulate real-world clinical inquiries. AI-generated TNM stages and treatment protocols were compared against AJCC 8th Edition guidelines and expert multidisciplinary tumor board decisions. The analysis focused on diagnostic accuracy, inter-rater agreement, and the impact of anatomical complexity on model performance. Results Both models demonstrated comparable proficiency in TNM staging accuracy (ChatGPT: 75.6%, Gemini: 75.0%), showing substantial agreement with expert standards (x 2 = 0.000, p = 1.000). However, a significant divergence was observed in treatment planning; Gemini achieved a 78.9% accuracy rate, significantly outperforming ChatGPT’s 71.7% (p=0.043 (x 2 = 4.114). Notably, ChatGPT’s staging performance was sensitive to tumor localization, with decreased precision in anatomically complex regions such as the oropharynx and paranasal sinuses (p = 0.034, Cramer’s V = 0.291). Conversely, Gemini demonstrated more robust spatial reasoning across different subsites. Conclusion While both LLMs provide reliable staging support, Gemini exhibits superior clinical reasoning in synthesizing multidimensional data into actionable treatment recommendations. However, a staging error rate of 25% remains a critical concern, potentially leading to inappropriate clinical pathways. These models should be viewed as auxiliary tools within an ‘augmented intelligence’ ecosystem, integrated with imaging and multidisciplinary inputs, rather than independent decision-makers. Strict expert supervision is mandatory to prevent subsite-specific errors from impacting surgical and oncological outcomes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Baklaci et al. (2026) studied this question.

synapsesocial.com/papers/6a0946ff8f2332546a459162https://doi.org/10.3389/fonc.2026.1828538
Ask AI
Helpful
Bookmark
Share
View Full Paper