INTRODUCTION: Patients increasingly use large language models (LLMs) for health information, yet the quality of fecal incontinence (FI) content from these models has not been evaluated with validated instruments. OBJECTIVE: Our objective was to characterize the quality, understandability, and actionability of information on FI generated by four LLMs. METHODS: Search data was extracted from Google Trends’ top five search queries related to FI. These queries were fed into four LLMs: ChatGPT version 4o mini (OpenAI), Perplexity version sonar (Perplexity AI), LLaMA version 3.2 (Meta AI), and Gemini version 2.0 flash (Google). Each query was asked in a separate instance to avoid bias in response. Responses were assessed using the following validated instruments by an expert panel of urogynecologists: DISCERN to assess the quality of written health care content and treatment choices, Patient Education Materials Assessment Tool (PEMAT) to assess understandability and actionability, and the Flesch–Kincaid Grade Level (FKGL) to analyze readability. Statistical analyses performed using R, version 4.4.2 (R Core Team, 2025). RESULTS: Response quality was moderate (median DISCERN 44 out of 80), understandability was high (median 77% PEMAT), and actionability was poor (median 40% PEMAT) across all four LLMs. DISCERN scores varied significantly (p<.0001), with Gemini scoring higher than ChatGPT (p<.001) and Perplexity scoring higher than ChatGPT (p<.001). Responses were at college level based on the FKGL score. Readability scores differed significantly (p<.001), with LLaMA generating text at a higher reading level than Perplexity (p<.001). CONCLUSIONS: LLM responses were understandable but lacked actionability and were above the recommended reading level for consumer health information. Perplexity and Gemini appeared to have better performance regarding FI queries. Other than Perplexity, none of the LLMs provided references for their responses. Given the evolving nature of this technology, continued evaluation of AI is essential before it can be accepted as an accurate and reliable source for patients. This research received no specific grant from any funding agency, commercial or not-for-profit sectors.Table 1
Hardy et al. (Fri,) studied this question.