This study investigated whether humans and generative Large Language Models (LLMs) exhibit similar performance in divergent ideation but diverge in convergent selection. To address the critical oversight in current AI creativity research, which predominantly focuses on generative output, this study introduces the original conceptual framework of ‘Selection Alignment’ and a ‘novel dual-phase experimental protocol.’ This research transcends traditional generation-centric evaluations to establish a new paradigm for assessing the evaluative stage of creativity. A controlled experiment involved 240 design professionals (120 idea generators, 120 independent selectors) and two LLM agents (GPT-4o, Gemini 1.5 Pro). Participants and LLMs responded to identical divergent prompts, including 10 Alternative Uses Task-style prompts and 10 design problems. Both humans and LLMs generated candidate idea pools, then performed convergent selection by choosing the top five items per prompt. Idea generation was evaluated based on Fluency, Flexibility, and Semantic Breadth. Selection outcomes were compared using top-5 overlap rates derived from semantic clustering. The results indicated near-parity in generation metrics, showing no statistically significant differences between human and AI outputs. However, a substantial divergence was observed in convergent selection: the mean human–AI top-5 overlap was 19.2% for Model-A and 22.4% for Model-B, both significantly below permutation-based chance levels (null mean overlap ≈ 35%). AI selections were strongly predicted by embedding- and probability-based metrics, while human choices were better predicted by context- and experience-based criteria, highlighting a fundamental mechanistic divide. This suggests that convergent selection amplifies human–AI divergence, carrying significant implications for designing co-creative interfaces that integrate human experience into AI’s selection mechanisms.
Jung et al. (Mon,) studied this question.