Large Language Models (LLMs) now occupy a prominent role in science, engineering, and higher education. Their capacity to generate step-wise solutions, conceptual explanations, and problem-solving pathways creates new opportunities—but also new risks—for Chemical Engineering learners. Despite widespread informal use, few empirical studies have evaluated LLM performance using a systematically designed dataset mapped directly to Bloom’s Taxonomy. This study evaluates the competency of ChatGPT in solving Chemical Engineering problems mapped to Bloom’s Taxonomy. A diverse dataset of undergraduate-level problems spanning six cognitive domains was used to assess the model’s reasoning across ascending levels of cognitive complexity. Each response was evaluated for accuracy and categorized into five error types. Although ChatGPT demonstrated considerable potential across a range of topics, the analysis also revealed important challenges and limitations that inform best practices for integrating LLMs into Chemical Engineering education. Results show significant differences in ChatGPT performance across Bloom levels, revealing three distinct tiers of capability. Strong performance was observed at lower cognitive levels (Remember–Apply), while substantial degradation occurred at Analyze, Evaluate, and specially Create. The findings provide a nuanced, empirically grounded understanding of current LLM capability limits, with practical recommendations for educators integrating LLMs into engineering curricula.
Shahid et al. (2026) studied this question.