PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026Journal of the American Medical Informatics Association0 citations

CPGPrompt: translating clinical guidelines into large language model-executable decision support

View Full Paper
RDRuiqi DengGMGreg S. MartinTWTony J. C. Wang

Key Points

  • The aim is to integrate clinical practice guidelines into AI systems for improved decision support in patient care.
  • Developed CPGPrompt, an auto-prompting system for converting guidelines into structured decision trees.
  • Generated synthetic vignettes across three health domains: headache, lower back pain, and prostate cancer.
  • Assessed system performance on binary specialty referral and multiclass pathway classification tasks.
  • Binary specialty referral classification achieved strong performance with F1 scores between 0.85 and 1.00 and perfect recall of 1.00.
  • Multiclass pathway assignment showed F1 scores of 0.47 for headache, 0.72 for lower back pain, and 0.77 for prostate cancer, reflecting domain-specific challenges.

Abstract

Abstract Objective Clinical practice guidelines (CPGs) provide evidence-based recommendations for patient care; however, integrating them into artificial intelligence (AI) remains challenging. Previous approaches, such as rule-based systems or black-box AI models, face significant limitations, including poor interpretability, inconsistent adherence to guidelines, and narrow domain applicability. To address this, we develop and validate CPGPrompt, an auto-prompting system that converts narrative clinical guidelines into large language models (LLMs). Materials and Methods Our framework translates CPGs into structured decision trees and utilizes an LLM to dynamically navigate them for patient case evaluation. Synthetic vignettes were generated across 3 domains—headache, lower back pain, and prostate cancer—and distributed into 4 categories to test different decision scenarios. System performance was assessed on both binary specialty referral decisions and fine-grained pathway classification tasks. Results The binary specialty referral classification achieved consistently strong performance across all domains (F1: 0.85-1.00), with high recall (1.00 ± 0.00). In contrast, multiclass pathway assignment showed reduced performance, with domain-specific variations: headache (F1: 0.47), lower back pain (F1: 0.72), and prostate cancer (F1: 0.77). Discussion Domain-specific performance differences reflected the structure of each guideline. The headache guideline highlighted challenges with negation handling. The lower back pain guideline required temporal reasoning. In contrast, prostate cancer pathways benefited from quantifiable laboratory tests, resulting in more reliable decision-making. Conclusion CPGPrompt demonstrates generalizability across diverse clinical domains while maintaining high sensitivity for referral decisions. Its transparent, auditable framework enables the systematic identification of failure modes and provides advantages over black-box AI approaches. However, persistent challenges with subjective clinical assessments indicate a need for targeted improvements and greater clinical robustness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Deng et al. (2026) studied this question.

synapsesocial.com/papers/69a286850a974eb0d3c01906https://doi.org/10.1093/jamia/ocag026
Ask AI
Helpful
Bookmark
Share
View Full Paper