Objective: Cardiovascular physiology underpins the understanding of blood pressure regulation, autonomic control, and hemodynamic mechanisms central to hypertension and cardiovascular disease. As generative artificial intelligence (AI) is increasingly adopted in medical education, its role in supporting rigorous assessment of cardiovascular knowledge remains uncertain. Evidence integrating learner performance, perceptions, and detailed psychometric behavior of AI-generated assessments in cardiovascular physiology is limited. Accordingly, this study aimed to compare AI-generated and expert-authored multiple-choice questions (MCQs) in cardiovascular physiology with respect to student performance, perceived quality, and item- and test-level psychometric properties. Design and method: In a blinded, cross-sectional study, 34 undergraduate medical and pharmacy students completed two content-equated MCQ examinations (20 items each): one generated using ChatGPT (OpenAI) and one authored by an experienced physiologist. Students were blinded to item origin and completed both assessments in randomized order. Performance was analyzed overall and across cognitive domains (knowledge, skill, competence). Students rated question quality and provided qualitative feedback. Item difficulty, discrimination, point-biserial correlations, non-functional distractors, and test-level reliability indices were evaluated. Multiple regression examined predictors of performance. Results: Students achieved higher overall scores on AI-generated items, driven by superior performance in knowledge and skill domains, with no difference in competence-level (integrative) reasoning. Medical students performed better on AI-generated items, whereas pharmacy students showed comparable performance across both papers. AI-generated questions were perceived as clearer and easier, while expert-authored items were more strongly associated with critical thinking and clinical realism. Item- and test-level psychometric indices were broadly comparable between examinations; however, AI-generated items exhibited a higher burden of non-functional distractors. No demographic or perception-based variables independently predicted performance. Conclusions: AI-generated MCQs can support assessment of foundational cardiovascular physiology knowledge with psychometric performance comparable to expert-written items. However, limitations in distractor quality and higher-order cognitive elicitation highlight the continued necessity of expert oversight. AI-assisted item generation is best positioned as a complementary tool to support cardiovascular education and training rather than a substitute for expert-authored assessments.
Alhamad et al. (Fri,) studied this question.