Evaluating AI-assisted tools in healthcare human–computer interaction (HCI) presents methodological challenges when practical constraints limit sample sizes. Standard pooled statistical analysis can then produce misleading results, including Simpson’s Paradox, where aggregate trends contradict patterns observed within subgroups. This paper introduces a conditional Gaussian model framework that models each experimental condition separately rather than pooling all observations. Through a within-subjects evaluation of an AI-assisted UI/UX design tool for medical software interfaces (n = 4 professional designers), we demonstrate how pooled analysis produced a misleading negative correlation between design time and IEC 62366 compliance (the medical device usability standard; pooled r=−0.76, p=0.029, n=8), even though every designer achieved both faster times and higher compliance with the AI tool. Within-condition correlations were non-significant and inconsistent in sign, confirming the pooled association as an aggregation artefact rather than a within-designer trade-off. The conditional analysis surfaces experience-indexed differences: the less UI-experienced designer showed the largest time reduction (up to 92%), while the two high-AI-experience designers showed the largest automated proxy-compliance gains (+25 to +29 percentage points). Sample standard deviations were also lower in the AI-assisted condition than in the traditional condition for both outcomes (time: 20.0→11.3 min; compliance: 10.6→7.6 percentage points); at n=4 per condition, however, this difference in variance can neither be confirmed nor falsified, and we make no inferential claim about variance compression. A follow-up phase (n = 3) that adapted the tool’s scaffolding to designer experience yielded a bidirectional response, with the two high-AI-experience designers further reducing time and the less UI-experienced designer engaging more deeply with the design output. Because all participants completed the traditional condition before the AI-assisted condition, the study is interpreted as a sequentially unbalanced exploratory comparison, not as a counterbalanced causal test of tool effectiveness. We provide guidelines for healthcare HCI researchers facing sample-size constraints endemic to specialised domains.
Firoz et al. (2026) studied this question.