Abstract Objective We examined whether GPT-4o, a widely used large language model (LLM), could produce age- and education-appropriate versions of complex pediatric traumatic brain injury case descriptions, while preserving clinical accuracy and emotional tone. Methods Five cases were adapted into four audience scenarios. Text complexity was assessed via Flesch–Kincaid (FKS), Gunning Fog, and SMOG indices. Clinical human experts rated text fidelity and emotional appropriateness on a 3-point scale. Results Original texts showed very high complexity (FKS 18.2–20.5), equivalent to 18–20 years of education. Adaptations for parents with high school education were often over-simplified (FKS 4.75–7.1), while versions for 12-year-olds were well-matched (FKS ~5–6). Texts for 8-year-olds had FKS scores of 4.0–6.8 (above grade 2–3 targets) and reduced fidelity (scores 1–2). Emotional tone was consistently rated appropriate across all audiences. Conclusion Clinicians may use LLMs to draft explanations, but must carefully review and tailor them.
García-Rudolph et al. (2026) studied this question.