Key points are not available for this paper at this time.
Background Artificial intelligence (AI) tools such as ChatGPT are increasingly being explored for clinical decision support, yet their role in geriatric medicine remains uncertain due to the complexity of multimorbidity and care planning. This study aimed to evaluate the clinical accuracy, completeness, and guideline alignment of ChatGPT's responses to common geriatric scenarios using standardized vignettes. Methodology Seven standardized vignettes representing common geriatric scenarios, namely, polypharmacy, falls, dementia, delirium, frailty, advance care planning, and urinary incontinence, were submitted to ChatGPT (GPT-5). Responses were evaluated by five independent consultant geriatricians using a standardized rubric across the following five domains: accuracy, completeness, guideline alignment, safety, and clarity (0-2 score per domain). Descriptive statistics summarized performance, and qualitative feedback was thematically analyzed. Inter-rater reliability was assessed using Krippendorff's alpha. Results ChatGPT scored the highest in clarity (66/70) and safety (63/70), with slightly lower performance in accuracy (59/70) and completeness (55/70). Guideline alignment was generally strong (61/70). Advance care planning received the highest domain scores; urinary incontinence scored the lowest. Krippendorff's alpha showed high inter-rater agreement (0.969). Reviewers identified key omissions, such as missing assessments or guideline-recommended tools, in multiple vignettes. Conclusions ChatGPT showed potential as a supportive tool in geriatric care, offering clear and generally safe responses aligned with guidelines. However, it lacked clinical depth and missed key elements in complex scenarios. AI tools such as ChatGPT should be used with caution, under expert oversight, and not as standalone decision makers in clinical practice.
Cassar et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: