Key points are not available for this paper at this time.
Background Large language models (LLMs) hold considerable potential in medical and health education; however, their reliability and interpretability in highly sensitive areas and in decision-making remain unclear. This study focuses on four publicly available LLMs and systematically evaluates their applicability in fertility preservation scenarios for breast cancer patients, thereby providing guidance for targeted use. Methods This study utilizes Google Trends to identify and filter information on topics related to fertility preservation for breast cancer patients, and analyses the dialogue outputs of models such as GPT-5.4 Thinking, Gemini 3.0, DeepSeek-V3.2, and Microsoft Copilot. To ensure consistency in responses and the fairness of LLM baseline performance, only one response is generated per query, and no responses are generated repeatedly; all dialogues are submitted to the four large language models using standardized prompts. The study found that 26 fertility-preserving response outputs in breast cancer patients exhibited varying patterns, revealing characteristics relevant to fertility-preserving treatments and decision-making for breast cancer patients. The study utilized reliability assessment tools, including DISCERN, EQIP (Evaluation of Information Quality to Patients), GQS (Global Quality Score) and JAMA (JAMA benchmark criteria), a comprehensive assessment based on six widely used readability metrics Automated Readability Index (ARI), Coleman–Liau Index (CLI), Flesch–Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Simple Measure of Gobbledygook (SMOG) with Flesch Reading Ease Score (FRES). Results The findings indicate that there are statistically significant differences in the reliability of various artificial intelligence programmes when it comes to providing highly sensitive, decision-intensive, multidisciplinary medical consultations regarding fertility preservation for breast cancer patients. The average intraclass correlation coefficients for all LLMs ranged from 0.715 to 0.978 (with all p -values 0.001). Microsoft Copilot demonstrates superior performance in terms of information reliability and structural quality, DISCERN 56.5 (49.25, 61), EQIP65.0 (51.25, 75.0), GQS3.0 (3.0, 4.0), JAMA 1.0 (1.0, 2.0), with a higher score than GPT-5.4 Thinking, Gemini 3.0 and DeepSeek-V3.2, the model is capable of providing more reliable information and better decision-making support. The responses generated by all LLMs are too complex for the general public and fail to meet the recommended reading comprehension standards for years 6 to 8; the writing standards of most outputs are equivalent to those of secondary school education, or the reading level required for legal documents. Conclusion This study reveals differences in the information provided by various LLMs regarding fertility preservation decisions for breast cancer patients, and recommends selecting a model suited to the specific clinical context; the Microsoft Copilot model demonstrated the best performance. Although LLMs demonstrate a certain degree of reliability when handling complex health enquiries, none have met the readability benchmark recommended for a year 6 reading level. Future research should focus on improving the reliability and readability of health information generated by LLMs to enhance comprehension among a wider audience.
Wang et al. (Fri,) studied this question.