Background: Pediatric urolithiasis is an increasingly important health concern, and affected children and their families require information that is both accurate and easily understandable. Artificial intelligence (AI)-powered chatbots have become widely used sources of health information; however, the readability, quality, and reliability of their outputs remain insufficiently evaluated. This study aimed to assess the effectiveness and reliability of AI chatbots in providing patient-oriented information on pediatric kidney stone disease and to identify factors influencing the quality and readability of their responses. Methods: Four AI chatbots (ChatGPT-5, Google Gemini, Claude 3 Opus, and DeepSEEK) were queried with 30 standardized questions related to pediatric kidney stones. Readability was evaluated using the Average Reading Level Consensus (ARLC), Automated Readability Index (ARI), and Simple Measure of Gobbledygook (SMOG). Response quality and reliability were asssessed using the Ensuring Quality Information for Patients (EQIP) tool and Modified DISCERN score. Statistical analyses included one-way analysis of variance ANOVA, Kruskal-Wallis tests, and appropriate post hoc comparisons. Results: Readability differed significantly among the chatbots. Google Gemini demonstrated the highest reading levels across all metrics (ARLC: 14.93, ARI: 16.2, and SMOG: 13.32), whereas ChatGPT, Claude, and DeepSEEK produced less complex test (p Conclusions: Substantial variability exists in the readability and reliability of AI-generated health information on pediatric urolithiasis. Although ChatGPT and Google Gemini provided more reliable information, Google Gemini’s responses were consistently more complex and less accessible. These findings emphasize the need for careful validation and language simplification of AI-generated content before its use in patient and caregiver education.
A Thu, study studied this question.