Objective: This study aimed to evaluate the quality of information provided by artificial intelligence models -ChatGPT, Gemini, and DeepSeek- when responding to common questions asked by patients diagnosed with bladder cancer. Material and Methods: A total of 30 frequently asked patient questions, obtained from reputable cancer information platforms and patient forums were submitted to each model under standardized conditions. The responses were independently evaluated by 2 experienced urologists using the DISCERN instrument and the Global Quality Scale (GQS). Only the initial responses were analyzed to reflect real-world patient behavior. Results: ChatGPT achieved the highest mean scores for both DISCERN (4.90±0.31) and GQS (4.83±0.38), followed by Gemini (DISCERN: 4.47±0.57, GQS: 4.53±0.57). DeepSeek demonstrated lower performance with mean DISCERN and GQS scores of 4.17±0.83 and 4.20±0.89, respectively. Pairwise comparisons revealed that ChatGPT significantly outperformed DeepSeek (p<0.05), whereas differences between ChatGPT and Gemini, and Gemini and DeepSeek, were not statistically significant. Conclusion: Artificial intelligence language models have potential as supplementary tools in patient education and counseling for bladder cancer. While ChatGPT and Gemini provided high-quality and reliable information, DeepSeek requires further refinement. Importantly, none of these models can substitute for professional medical consultation, and their role should remain supportive rather than directive in clinical decision-making.
Demir et al. (Thu,) studied this question.