PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 28, 2026Cureus0 citationsOpen Access

Detecting Citation Veneer in Guideline-Grounded Clinical and Public Health Large Language Model (LLM) Outputs: A Clinical Evidence Audit Grid Using CENHOV Tags

YHYuusuke Harada

Key Points

  • The aim is to identify and measure errors in outputs from large language models used for clinical guidelines.
  • Developed the Clinical Evidence Audit Grid (CEAG) for monitoring outputs.
  • Evaluated CEAG on a synthetic benchmark with 120 scripted questions and a 12-case public check.
  • Used binary signals to assess correctness and citation support in LLM outputs.
  • In the synthetic benchmark, evidential support was 93.3% but correctness dropped to 73.3% with 25.8% citation veneer.
  • In the public check, RUSH generated 66.7% citation veneer, while VERIFY and POSTER produced none.

Abstract

Large language model (LLM) systems are increasingly used to summarize clinical guidelines, draft patient-facing education, and answer evidence-linked clinical or public health questions. We define citation veneer as an observable audit state in which an output presents citation cues or apparently supportive source information while still containing incorrect, incomplete, or materially unsupported content. We develop and evaluate the Clinical Evidence Audit Grid (CEAG), an analyst-facing monitoring surface that encodes six binary response-quality signals into a fixed-length plain-text audit tag (CENHOV): clinical correctness (C), evidential support (E), numeric concordance (N), hedging (H), stylistic ornamentation (O), and a derived CEAG-V supported-error citation veneer marker (V=1 when E=1 and C=0). CEAG was evaluated on a controlled benchmark (synthetic, n=120 questions per condition) and on a 12-case public check built from PubHealth-style health-claim examples and public medical or public-health sources. The controlled benchmark used scripted response regimes rather than a live LLM/API benchmark; therefore, model identity, sampling temperature, and seed are not applicable to the main controlled analysis. In the synthetic benchmark, the RUSH regime retained high evidential-support rates (93.3%, 95% CI 87.4%-96.6%) while dropping to 73.3% correctness (95% CI 64.8%-80.4%) and producing 25.8% CEAG-V veneer (95% CI 18.8%-34.3%). In the 12-case public check, RUSH produced 66.7% CEAG-V veneer (95% CI 39.1%-86.2%), whereas VERIFY and POSTER produced none; because n=12, these public-case rates are descriptive pilot data only. CEAG is not proposed as a patient-facing display or a replacement for conventional charts. It is a clinical informatics monitoring aid for reviewers who need to detect apparently well-sourced hallucinations, compare response regimes, and separate stylistic change from substantive trust change.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yuusuke Harada (2026) studied this question.

synapsesocial.com/papers/69f04e5b727298f751e72440https://doi.org/10.7759/cureus.107729
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Benchmarking Large Language Models in Retrieval-Augmented Generation2024 · 373 citations
  2. 2Toward a responsible future: recommendations for AI-enabled clinical decision support2024 · 220 citations
  3. 3Do no harm: a roadmap for responsible machine learning for health care2019 · 1,208 citations
  4. 4FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation2023 · 331 citations
  5. 5Key challenges for delivering clinical impact with artificial intelligence2019 · 2,957 citations