PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 6, 2026European Heart Journal0 citations

How do current artificial intelligence chat-based large language models perform as a patient educational tool for atrial fibrillation ablation?

View Full Paper
CHC G HongLDL DagherTBT Baykaner

Key Result

Four large language models evaluated for atrial fibrillation ablation patient education failed to meet the recommended 8th-grade reading level (mean readability 12.3) and were only reasonably accurate.

Key Points

  • This research evaluates how well AI language models can provide understandable and accurate information about atrial fibrillation ablation.
  • Analyzed responses to 25 prompts from 4 AI language models regarding AF ablation.
  • Used readability tests: Flesh-Kincaid and SMOG to assess text difficulty and grade level.
  • Evaluated accuracy using a modified Likert-like scale by experienced electrophysiologists.
  • Overall mean word count was 254.9, with WikiCardio having the longest average.
  • Readability scores averaged at 12.3 by FK and 11.7 by SMOG, indicating high difficulty levels.
  • Only 2 out of 100 responses achieved an 8th-grade reading level, suggesting poor clarity for patients.

Study Design

Type

Cross-Sectional

Blinding

Blinded

Structured PICO

Do current artificial intelligence chat-based large language models provide readable and accurate patient educational information for atrial fibrillation ablation?

P
Population
25 atrial fibrillation ablation prompts sourced from academic healthcare organizations' patient webpages
I
Intervention
4 free large language models (ChatGPT4o mini, Google Gemini 1.5 Flash, Microsoft CoPilot, and WikiCardio Virtual Assistance)
O
Outcome
Readability (Flesh-Kincaid Grade Level and SMOG scores) and accuracy (graded by 3 blinded electrophysiologists on a 1-5 scale)

Current LLMs provide information on AF ablation that is too difficult for average patients to read and only reasonably accurate, highlighting the need for further refinement before clinical use.

Abstract

Abstract Background Atrial fibrillation (AF) affects nearly 46.3 million individuals globally (1). Recent guidelines recommend catheter ablation as first-line therapy in appropriate patients, but centralized patient educational materials on ablation are lacking (2). Chat-based large language models (LLMs) have become popular as informational resources. Whether LLMs can serve as medically accurate and comprehensible patient educational tools for AF ablation has not been investigated. Purpose We assessed the readability and accuracy of responses from 4 LLMs to potential patient questions ("prompts") regarding AF ablation. Methods Twenty-five AF ablation prompts were sourced from academic healthcare organizations’ patient webpages. Prompts were fed to 4 free LLMs: ChatGPT4o mini (CG), Google Gemini 1.5 Flash (GG), Microsoft CoPilot (MCP), and WikiCardio Virtual Assistance (World Heart Federation) (WK) (Fig. 1). Flesh-Kincaid (FK) Grade Level and SMOG (SimpleMeasure of Gobbledygook) readability scores were used to assess responses for word count (WoC), difficulty, and grade-level, as ≤8th-grade reading level is recommended for patient information. Responses were graded by 3 experienced electrophysiologists using a modified Likert-like scale: "5-highly accurate ("A", 90% accurate); 4-mostly accurate ("B", 80% accurate); 3-reasonably accurate/passable ("C", 70% accurate); 2-mostly inaccurate ("D", 60% accurate); 1-highly inaccurate/misleading ("F", ≥50% inaccurate)". All evaluations were blinded. Results Overall mean WoC was 254.9; WK had the longest mean WoC (304±93.1 vs CG 286.7±65.6, MCP 247.4±82, GG 181.4±79.5) (Fig. 2a). Overall mean readability scores were 12.3 by FK and 11.7 by SMOG. WK had the highest mean readability scores 14.2±0.9 by FK (Professional; College), 13.3±0.6 by SMOG (Very Difficult; College level Entry; CG, GG, and MCP were at 12th-grade level (Difficult) by FK CG 11.7±1.3 vs GG 11.6±1.8 vs MCP 11.7±1.3 and 11th grade (Fairly Difficult) by SMOG CG 11.1±1.0 vs GG 11.3±1.5 vs MCP 11.2±1.1 (Fig. 2b-c). Only 2 out of 100 total LLM responses scored at an 8th-grade reading level. Overall content evaluation scores were 3.4 mean and 4 median; CG was the highest (3.5 mean, 4 median), GG was the lowest (3.3 mean, 3 median) (Fig 2d). GG had 16% "highly inaccurate/misleading" responses; WK was unable to respond to 3 (12%) prompts. MCP had 82% follow-up questions asked back to the user; 100% of responses from CG and WK emphasized further discussion with a healthcare team. Conclusion In their current states, LLMs are not readable nor reliable sources for patient information regarding AF ablation. All 4 LLMs failed to meet the recommended 8th-grade reading level and were only reasonably accurate for content; only 2 models advised discussion with a healthcare team. These findings suggest the need to further refine LLM algorithms and potentially involve patient advocates in their continual development as potential educational tools.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hong et al. (2025) conducted a cross-sectional in Atrial fibrillation. Large language models (ChatGPT4o mini, Google Gemini 1.5 Flash, Microsoft CoPilot, WikiCardio Virtual Assistance) was evaluated on Readability (Flesh-Kincaid Grade Level and SMOG) and accuracy (modified Likert-like scale). Four large language models evaluated for atrial fibrillation ablation patient education failed to meet the recommended 8th-grade reading level (mean readability 12.3) and were only reasonably accurate.

synapsesocial.com/papers/698585fe8f7c464f23009c78https://doi.org/10.1093/eurheartj/ehaf784.4502
Ask AI
Helpful
Bookmark
Share
View Full Paper