Efficient, low-workload conversational support is critical when untrained bystanders must deliver urgent first aid. We compared a role-specialized multi-agent system (MAS-EC) with a state-of-the-art large-language model (GPT-4o) across cardiopulmonary resuscitation (CPR) and epinephrine-auto-injector scenarios in a fully counter-balanced, repeated-measures experiment involving 13 participants. Behavioral efficiency (clarification-query count), subjective workload (NASA-TLX), and usability (SASSI) were recorded. Repeated-measures ANOVA showed that MAS-EC cut queries by 56 % ( F = 18.84, p = .001) and reduced overall NASA-TLX scores ( F = 34.95, p < .001). An AI × Task interaction indicated GPT overhead was negligible for CPR but tripled for the more cognitively demanding epinephrine scenario ( F = 20.09, p < .001), whereas SASSI ratings did not differ. Poisson GLMM confirmed efficiency findings despite non-normal count data. The study demonstrates that decentralizing expertise across cooperative agents simultaneously improves interaction efficiency and lowers cognitive load without sacrificing perceived usability. Limitations include a modest sample, text-only simulations, and a single MAS-EC configuration; future work should test multimodal conditions and additional ensemble strategies.
Walji et al. (2026) studied this question.