Multi-agent collaboration now underpins critical systems-from warehouse swarms to LLM-based agentic workflows-yet collaboration is never free. It competes with hard resource budgets (bandwidth, latency, energy, compute) and soft frictions (partial observability, strategic coupling, and non-stationarity). While prior surveys largely catalog method families, they often fail to address the practitioner's fundamental decision problem: determining when collaboration justifies its cost and how to design mechanisms that survive real-world constraints. We present a decision-oriented survey structured around four pivotal research questions. RQ1 (Communication) investigates the value of interaction under explicit budgets, identifying when shared representations, topology design, or temporal sparsity can substitute for costly messaging. RQ2 (Credit Assignment) focuses on aligning local learning with system-level objectives, synthesizing mechanisms-from counterfactual baselines and Shapley values to structured critics-that ensure faithful attribution under heterogeneity. RQ3 (Scalability) examines the boundaries of collaboration, analyzing saturation effects, cascading failures, and the onset of negative collaboration as populations and task densities scale. RQ4 (Evaluation) proposes decision-aligned protocols that replace single-point metrics with efficiency frontiers and deployability gates, explicitly measuring alignment and trustworthiness alongside capability. Synthesizing empirical regularities and decision rules across these dimensions, we outline a compact reporting standard and future trends toward adaptive orchestration and decentralized ecosystems. Our goal is to transform a fragmented evidence base into defensible guidance for designing and deploying collaborative AI systems at scale.
Cheng et al. (Tue,) studied this question.