Abstract Cognitive Diagnostic Models (CDMs) provide fine‐grained diagnostic feedback, but their central component—the Q‐matrix—remains costly and labor‐intensive to construct. This study explores the automated generation of Q‐matrices using general‐purpose AI, including ChatGPT‐4o, Gemini‐2.5‐pro, and Claude‐sonnet‐4. We evaluated two prompting strategies (all‐at‐once and one‐by‐one) across TIMSS 2007, TIMSS 2011, and PISA 2012 mathematics assessments. Results show that AI‐generated Q‐matrices approximate human baselines with competitive model fitting performance (AIC, BIC, log‐likelihood, and SRMSR) and acceptable classification discrepancies. While AI predictions for larger and more complicated assessments (TIMSS 07 and 11) were generally sparser than human‐generated Q‐matrices, they still achieved equal or better fit statistics under most CDMs. In contrast, for the smaller and less complicated PISA 2012 assessment, AI‐generated Q‐matrices matched human density and fitting quality. Importantly, chatbot‐human matching accuracy remained high across models, with Gemini benefiting from all‐at‐once prompting, ChatGPT‐4o maintaining stable performance under both strategies, and Claude showing sensitivity to prompt structure. These findings highlight both the promise and current limitations of automated Q‐matrix generation, underscoring opportunities for integrating LLMs into scalable diagnostic assessment practices.
Xue et al. (Thu,) studied this question.