PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 20, 2026Journal of Diabetes Science and Technology0 citations

Evaluating the Accuracy, Quality, and Reproducibility of AI Chatbot-Generated Diabetes Self-Management Education and Support Plans Aligned With ADCES7

View Full Paper
TYTeng‐Hung YuHHHui‐Chun HsuYLYau‐Jiunn Lee

Key Points

  • This study assesses the quality, accuracy, and reproducibility of diabetes self-management plans generated by an AI chatbot.
  • Evaluated eleven virtual patient profiles across five key diabetes self-management time points.
  • Used structured prompts with retrieval-augmented generation to generate plans assessed with DSMES Process Evaluation Checklist.
  • Engaged ten certified diabetes educators to score the plans and examined internal consistency, inter-rater reliability, and reproducibility.
  • Mean domain scores ranged from 29.0 to 42.7, with high scores in clarity (85.1%) and goal-directedness (81.1%).
  • Moderate performance in accuracy (74.9%) and feasibility (75.3%) observed, but limitations in patient engagement noted.
  • Reliability was high with a Cronbach's α of 0.889 and reproducibility moderate with intra-assay coefficients of 0.04 to 0.15.

Abstract

BACKGROUND: The rapid development of artificial intelligence, particularly large language models (LLMs) such as ChatGPT, Gemini, and Claude, offers new opportunities to scale and personalize diabetes self-management education and support (DSMES). This study evaluated the quality, accuracy, and reproducibility of DSMES plans generated by ChatGPT-5 using the ADCES7 Self-Care Behaviors™ framework. METHODS: Eleven virtual patient profiles representing five key DSMES time points (diagnosis, annual review, complicating factors, transitions, and ongoing care) were analyzed using structured prompts with retrieval-augmented generation. Generated plans were assessed using a DSMES Process Evaluation Checklist integrating IMPACTS, QAMAI, PDQI-9, and Donabedian's frameworks across ten domains. Content validity was confirmed by five experts (S-CVI: 0.982 for appropriateness; 0.964 for clarity). Ten certified diabetes educators (mean age 47.9 years; mean experience 14 years) independently scored the plans. Internal consistency, inter-rater reliability, and reproducibility were examined. RESULTS: Mean domain scores ranged from 29.0 ± 2.0 (patient engagement; 58.0%) to 42.7 ± 2.0 (cultural and linguistic appropriateness; 85.5%). ChatGPT-5 demonstrated strengths in clarity (85.1%) and goal-directedness (81.1%), generating evidence-based plans within three to five minutes, but showed limitations in patient engagement. Moderate performance was observed for accuracy (74.9%), relevance (78.7%), individualization (79.6%), and feasibility (75.3%). Reliability was high (Cronbach's α = 0.889; intra-class correlation coefficient = 0.837). Reproducibility was moderate, with intra-assay coefficients of 0.04 to 0.15 and inter-assay coefficients of 0.10 to 0.27. CONCLUSIONS: The DSMES process evaluation checklist is valid and reliable. ChatGPT-5 shows promise as a scalable DSMES decision-support tool, though improvements in personalization, patient engagement, and reproducibility are needed before broader implementation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yu et al. (2026) studied this question.

synapsesocial.com/papers/6a0d4f92f03e14405aa9af25https://doi.org/10.1177/19322968261447304
Ask AI
Helpful
Bookmark
Share
View Full Paper