PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026Acta cardiologica. Supplementum0 citations

Evaluating ChatGPT’s adherence to evidence-based heart failure guidelines: a comparative analysis using the 2023 ESC and 2022 ACC/AHA/HFSA recommendations

View Full Paper
MAMohammed Jassim AlnuwaysirAAAbdullah Shaker AljamaAGAbdulkarim Kassim Abdulgalil Galib

Key Points

  • Evaluate ChatGPT-5's accuracy and guideline adherence in heart failure management using established clinical recommendations.
  • Presented 38 anonymized heart failure clinical vignettes to ChatGPT-5.
  • Two board-certified cardiologists graded responses for guideline concordance on a 4-point scale.
  • Calculated descriptive statistics for performance and inter-rater agreement.
  • 20 (53%) responses were fully concordant with guidelines.
  • 6 (16%) responses unsafe or harmful, particularly in complex therapy cases.
  • Inter-rater agreement among cardiologists was high.

Abstract

BACKGROUND: Heart failure (HF) remains a major cause of morbidity and mortality worldwide. Large language models (LLMs) such as ChatGPT are emerging as potential clinical decision support tools, but their adherence to specialty guidelines is not well characterised. OBJECTIVES: To evaluate the accuracy and guideline concordance of ChatGPT-5 in managing real-world HF scenarios compared with the 2023 European Society of Cardiology (ESC) and 2022 American College of Cardiology (ACC)/American Heart Association (AHA)/Heart Failure Society of America (HFSA) recommendations. METHODS: Thirty-eight anonymised HF clinical vignettes spanning reduced, mildly reduced, and preserved ejection fraction phenotypes and varied New York Heart Association (NYHA) classes were presented to ChatGPT-5. Two board-certified cardiologists independently graded each response for concordance with guideline recommendations using a 4-point scale (3 = fully concordant, 2 = partially concordant, 1 = discordant, 0 = unsafe/harmful). Discrepancies were adjudicated by a third reviewer. Descriptive statistics summarised performance and inter-rater agreement. RESULTS: Of the 38 responses, 20 (53%) were fully concordant, 4 (11%) partially concordant, 8 (21%) discordant, and 6 (16%) unsafe/harmful. Most inaccuracies involved vague drug titration guidance, incomplete device therapy recommendations, or omission of guideline-directed medical therapy (GDMT). Unsafe suggestions occurred in complex device or advanced therapy decisions. Inter-rater agreement was high. CONCLUSIONS: ChatGPT-5 showed moderate concordance with ESC and ACC/AHA/HFSA HF guidelines, indicating potential value as a tool for knowledge synthesis and preliminary clinical support. However, its outputs require expert validation, and safe clinical integration will depend on future models incorporating guideline-based frameworks, real-time data, and rigorous physician oversight.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alnuwaysir et al. (2026) studied this question.

synapsesocial.com/papers/69fecf16b9154b0b828762a3https://doi.org/10.1080/00015385.2026.2668803
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 12022 AHA/ACC/HFSA Guideline for the Management of Heart Failure: A Report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines2022 · 2,246 citations
  2. 22021 ESC Guidelines for the Diagnosis and Treatment of Acute and Chronic Heart Failure2022 · 2,216 citations
  3. 3Artificial Intelligence and Heart Failure: A State-of-the-Art Review2023 · 91 citations
  4. 4Assessing the Health Education Needs of Heart Failure Patients in Saudi Arabia2024 · 1 citations
  5. 5ChatGPT Performance Deteriorated in Patients with Comorbidities When Providing Cardiological Therapeutic Consultations2025 · 3 citations